Compare commits

...
957 Commits
Author SHA1 Message Date
Ryan Houdek 93c428ae53 Docs: Update for release FEX-2508.1 2025-08-05 19:51:54 -07:00
Alyssa Rosenzweig 5836309525 unittests: add blake3 test
this provokes RA spilling and hit an assertion fail on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-05 19:50:53 -07:00
Alyssa Rosenzweig 43092ce48b RegisterAllocationPass: fix SRA spilling corner
I hate this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-05 19:50:48 -07:00
Alyssa Rosenzweig e5d51a20b2 RegisterAllocationPass: simplify next-use logic
I doubt this will fix the regression but it might make it easier to identify.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-05 19:50:43 -07:00
Ryan Houdek 59659a7184 Docs: Update for release FEX-2508 2025-08-01 18:24:33 -07:00
Ryan Houdek e6de17e72e Merge pull request #4752 from alyssarosenzweig/opt/defer-next-use-analysis-2
RegisterAllocationPass: defer next-use analysis
2025-08-01 16:38:31 -07:00
Ryan Houdek 0457bdc7ce Merge pull request #4751 from lioncash/perm
ASIMDOps: Remove unused permute overloads
2025-08-01 12:45:50 -07:00
Ryan Houdek 7e54c2735c Merge pull request #4750 from lioncash/op6
ASIMDOps: Move remaining base opcodes into implementing function
2025-08-01 12:45:24 -07:00
Alyssa Rosenzweig 7bcc58687f IR: remove a bunch of unused atomic ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:44 -04:00
Alyssa Rosenzweig 6a03df7b8d RedundantFlagCalculationElimination: drop dead syscall/atomic opts
I don't think these are worth it, and also currently they don't trigger ever.

n=100:
Difference at 95.0% confidence
	-0.00245961 +/- 0.00139573
	-0.524468% +/- 0.297615%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:37 -04:00
Alyssa Rosenzweig 7ee5a8065c RegisterAllocationPass: defer next-use analysis
This is expensive and only needed for spilling, so only do it for spilling. This
complicates the RA a bit but speeds us up on average since most blocks
don't spill. Total results of this change (including the prep commits that
slowed things down temporarily):

Difference at 95.0% confidence
	-1.71952% +/- 0.455996%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig 2f2353765f RegisterAllocationPass: use kill bits
this is a lot lighter weight than next uses for the same purpose.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig 8b4c2f5093 RegisterAllocationPass: consider AnySpilled at start of iteration
if we spill for SRA, we don't need/want to execute this code path. this will be
load bearing by the end of this series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig d80662a9ba RegisterAllocationPass: ignore kill bit in SRA
needed for the backwards pass internally due to ordering. a little awkward but
shrug.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 12:11:53 -04:00
Alyssa Rosenzweig 68c5c72dbc IR: model kill bits
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 12:04:50 -04:00
Lioncache b460cbc91a ASIMDOps: Remove unused permute overloads
These aren't used at all and don't really provide anything that the existing
non-templated overloads can't.
2025-08-01 11:51:55 -04:00
Lioncache f7939078c2 ASIMDOps: Remove unnecessary ins() overload
There's no difference between this and the non-templated version, in fact,
this variant wasn't even used at all, so we can just remove it.
2025-08-01 11:24:44 -04:00
Lioncache e103af3b93 ASIMDOps: Move base opcode into ASIMDScalarCopy()
Now we have no more duplicated opcodes in ASIMDOps.
2025-08-01 11:13:58 -04:00
Lioncache 45bec06d5b ASIMDOps: Move base opcode into ASIMDFloatConvBetweenInt()
Deduplicates the second last remaining instruction category in ASIMDOps
2025-08-01 10:49:52 -04:00
Alyssa Rosenzweig c82efe7621 RegisterAllocationPass: set AnySpilled less
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 10:39:11 -04:00
Ryan Houdek ba85dbf526 Merge pull request #4749 from lioncash/op5
ASIMDOps: Move most remaining base opcodes into their implementing functions
2025-07-31 12:21:31 -07:00
Lioncache e624e87b29 ASIMDOps: Move base opcode into ASIMDExtract()
Minor deduplication of the base opcode.
2025-07-31 10:44:06 -04:00
Lioncache 6ff8d0953f ASIMDOps: Move base opcode into ASIMDPermute()
Gets rid of the need to respecify the base opcode in every instruction implementation
2025-07-31 10:43:17 -04:00
Lioncache 9f2fcdf0b3 ASIMDOps: Move base opcode into Crypto2RegSHA()/Crypto3RegSHA()
Removes the need to specify the base opcode multiple times.
2025-07-31 10:23:38 -04:00
Lioncache cecb9bfe32 ASIMDOps: Move base opcode into CryptoAES()
Minor deduplication.
2025-07-31 10:15:15 -04:00
Lioncache 6a989d844a ASIMDOps: Move base opcode into ASIMDTable()
Gets rid of duplication of the opcode in all instruction implementations.
2025-07-31 10:10:35 -04:00
Ryan Houdek 318b3115a8 Merge pull request #4748 from lioncash/op4
ASIMDOps: Constrain more instructions with IsQOrDRegister
2025-07-30 16:22:14 -07:00
Lioncache 017952e89e ASIMDOps: Constrain more instructions with IsQOrDRegister
We had quite a few instructions that we weren't constraining with this,
now the instruction itself will show up in the failure output more immediately
instead of failing on the internal implementation function if anything is
incorrectly passed through.
2025-07-30 19:03:33 -04:00
LC b0c74e1b66 Merge pull request #4747 from Sonicadvance1/cpuid
CPUID: Update documentation comments
2025-07-30 18:31:54 -04:00
Ryan Houdek a190086504 Merge pull request #4746 from lioncash/op3
ASIMDOps: Move base opcode into ASIMDModifiedImm()/ASIMDShiftByImm()
2025-07-30 15:30:37 -07:00
Ryan Houdek 84704d1cc2 CPUID: Update documentation comments
Additional reserved bits have set uses now.
Additionally set the cpuid bit for bus-lock-detect, because FEX
definitely detects bus-locks.
2025-07-30 15:13:08 -07:00
Lioncache b1ef8bbc5a ASIMDOps: Move base opcode into ASIMDModifiedImm() 2025-07-30 18:03:43 -04:00
Lioncache f19a6343dd ASIMDOps: Move base opcode into ASIMDShiftByImm()
Avoids needing to duplicate the opcode all over the place
2025-07-30 18:03:36 -04:00
Ryan Houdek 23d0d7d4a4 Merge pull request #4745 from lioncash/op2
ASIMDOps: Move base opcode into ASIMD2RegMisc()
2025-07-30 13:50:04 -07:00
Lioncache ed7bc28c21 ASIMDOps: Simplify size conditionals in several 2-reg misc instructions
In many of these, the long ternaries are equivalent to a subtraction by 1,
which is much more efficient
2025-07-30 16:31:33 -04:00
Lioncache 05dcb160bd ASIMDOps: Move base opcode into ASIMD2RegMisc()
Avoids the need to specify the opcode repeatedly, making implementations less noisy.
2025-07-30 16:25:24 -04:00
Ryan Houdek e49155ff90 Merge pull request #4744 from lioncash/op
ASIMDOps: Move base opcode into implementation function for some categories
2025-07-30 12:46:07 -07:00
Ryan Houdek 265525e472 Merge pull request #4740 from Sonicadvance1/multiple_segments
FEXCore/Frontend: Ensure multiple prefix bytes work
2025-07-30 12:43:51 -07:00
Ryan Houdek d4dcbfa90b FEXCore/Frontend: Ensure multiple prefix bytes work
Only the last prefix byte is retained when multiple are set. We were
accidentally generating a mask.

Additionally with 64-bit code, the legacy segment prefixes don't
overwrite if FS or GS have been set. So no weird behaviour where FS/GS
is set, a legacy prefix is used for padding, and then it "ignores" a bad
prefix by ignoring only the latest one.
2025-07-30 12:30:13 -07:00
Ryan Houdek 6724673482 Merge pull request #4739 from Sonicadvance1/remove_check
Frontend: Remove arbitrary check
2025-07-30 12:26:13 -07:00
Ryan Houdek d04f75df29 Frontend: Remove arbitrary check
REX prefix isn't even encoded in to the instruction tables if a 32-bit
process is running. Just remove this.
2025-07-30 11:53:15 -07:00
Ryan Houdek a5260f4233 Merge pull request #4738 from bylaws/peggle
Implement inline SMC handling for linux FEX
2025-07-30 11:48:28 -07:00
Ryan Houdek 369ca5cb72 Merge pull request #4719 from Sonicadvance1/runtime_mode_switch_take2
Runtime mode switch take 2
2025-07-30 11:47:54 -07:00
Lioncache a3e02cbafa ASIMDOps: Simplify conditionals in saddlv/uaddlv
Really all these size conversions are emulating is a subtraction by 1.
Also we can drop in an assert that was missed in uaddlv
2025-07-30 11:21:06 -04:00
Lioncache 4cca2f6915 ASIMDOps: Move base opcode into ASIMDAcrossLanes() 2025-07-30 11:07:08 -04:00
Lioncache 64ca48c4b1 ASIMDOps: Move base opcode into ASIMD3Different()
Deduplicates the open-coded base opcode in the instruction implementations.
2025-07-30 10:59:11 -04:00
Lioncache 74a897271f ASIMDOps: Move base opcode into ASIMD3Same()
Moves the base opcode into the actual implementation, so that we
aren't open-coding it into every relevant instruction function.
2025-07-30 10:40:17 -04:00
Tony Wasserka c6733a6eec Merge pull request #4731 from Sonicadvance1/armtifacts
github: Upload armtifacts
2025-07-30 10:22:49 +02:00
Tony Wasserka f7e99678b4 Merge pull request #4730 from Sonicadvance1/fix_thunk_functional_path
unittests: Fixes thunk unittest path
2025-07-30 10:20:46 +02:00
Tony Wasserka 976b68ffec Merge pull request #4713 from Sonicadvance1/free_the_stats
Profiler: Decouple profile stats from the profiler option
2025-07-30 10:19:07 +02:00
Billy Laws 7b656fb009 SyscallsSMCTracking: Support inline SMC 2025-07-30 00:13:27 +01:00
Billy Laws 20331d52c3 LinuxSyscalls: Always reconstruct RIP and EFLAGS when spilling from the JIT 2025-07-30 00:13:27 +01:00
Billy Laws 2556acb82d TestHarnessRunner: Avoid frontend SMC handling 2025-07-30 00:13:27 +01:00
Ryan Houdek 248f0948b8 github: Upload armtifacts 2025-07-29 12:02:57 -07:00
Ryan Houdek 6c12db500e InstCountCI: Update for segment changes 2025-07-29 12:02:38 -07:00
Ryan Houdek 2a8f4dbbb6 Arm64EC: Update for GDT 2025-07-29 12:02:38 -07:00
Ryan Houdek 91598d5178 WOW64: Update for GDT 2025-07-29 12:02:37 -07:00
Ryan Houdek a6bb9739d4 OpcodeDispatcher: Initial support for runtime long-mode switch
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.

Adds some asserts since currently it is unexpected if the configuration
changes at runtime.

This is fairly straightforward for an initial setup but isn't fully
fleshed out.

Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.

Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2025-07-29 12:02:37 -07:00
Ryan Houdek aa871c797b FEXCore: Accurately store segment descriptors
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.

In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.

Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
2025-07-29 12:02:37 -07:00
Ryan Houdek 153d20ca59 Rename SHMStats 2025-07-29 12:02:19 -07:00
Ryan Houdek 03b04771be TestHarnessRunner: Become a real thread
Stop being so special as a host runner.
2025-07-29 12:02:19 -07:00
Ryan Houdek 392fa62dae Profiler: Decouple profile stats from the profiler option
This option is free and only enabled if the config option is set. Enable
it always at build time so that users can pick it up without enabling
the full gpuviz/tracy paths.
2025-07-29 12:02:19 -07:00
Ryan Houdek af3918491c unittests: Fixes thunk unittest path
This hasn't been run in CI for awhile so this was missed when I was
testing it. I need to see what I can do to get this going in CI again.
2025-07-29 11:46:30 -07:00
LC c0762bbd82 Merge pull request #4737 from Sonicadvance1/i_dislike_tuple_13
FileManagement: Remove pair usage from GetEmulatedFDPath
2025-07-29 10:48:38 -04:00
LC b68272b413 Merge pull request #4733 from Sonicadvance1/i_dislike_tuple_9
OpcodeDispatcher: Remove pair usage from DecodeNZCVCondition
2025-07-29 10:46:08 -04:00
LC 90b9414b5f Merge pull request #4732 from Sonicadvance1/i_dislike_tuple_8
FEXCore: Remove unused refcount_shared_mutex
2025-07-29 10:43:25 -04:00
LC 5db494d4f7 Merge pull request #4736 from Sonicadvance1/i_dislike_tuple_12
x87StackOptimizationPass: Removes pair usage
2025-07-29 10:42:55 -04:00
Tony Wasserka d073066293 Merge pull request #4722 from Sonicadvance1/fix_4124
CMake: Work around QCom disabling SVE in their chips
2025-07-29 10:54:18 +02:00
LC e80270ae43 Merge pull request #4735 from Sonicadvance1/i_dislike_tuple_11
64BitAllocator: Removes pair usage in allocator
2025-07-28 21:41:36 -04:00
LC cd56b85e88 Merge pull request #4734 from Sonicadvance1/i_dislike_tuple_10
Arm64: Remove pair usage in 128-bit loader
2025-07-28 21:40:57 -04:00
Ryan Houdek 64fbf55bdd FileManagement: Remove pair usage from GetEmulatedFDPath
NFC
2025-07-28 16:21:49 -07:00
Ryan Houdek 14c1ee10b6 x87StackOptimizationPass: Removes pair usage
NFC
2025-07-28 16:14:49 -07:00
Ryan Houdek 96ae671738 64BitAllocator: Removes pair usage in allocator
NFC
2025-07-28 16:07:07 -07:00
Ryan Houdek 16552d1194 Arm64: Remove pair usage in 128-bit loader
NFC
2025-07-28 15:58:11 -07:00
Ryan Houdek 6470c98ee2 OpcodeDispatcher: Remove pair usage from DecodeNZCVCondition
NFC
2025-07-28 15:50:31 -07:00
Ryan Houdek 126c4efa0e FEXCore: Remove unused refcount_shared_mutex 2025-07-28 15:43:44 -07:00
Ryan Houdek f6fd9e18d9 Merge pull request #4727 from bylaws/meopd
Windows: Lock invalidation tracking for the entire duration of memory ops
2025-07-28 11:59:23 -07:00
Ryan Houdek 1f554867e0 CMake: Work around QCom disabling SVE in their chips
Fixes #4124

clang feature checking can't check beyond MIDR, so compiling for a
specific cortex version means compiling SVE on these CPUs that disabled
them.

Just detect the particular situation in-which SVE isn't inside
proc/cpuinfo and is one of the snapdragon cores that are supposed to
support SVE. Then compile for Cortex-a78 instead.
2025-07-28 11:57:02 -07:00
Ryan Houdek d226331ccc Merge pull request #4726 from bylaws/ijwidjn
WOW64: Wrap BTCpuSimulate to ensure correct unwinding
2025-07-28 11:50:34 -07:00
Ryan Houdek 123f8c93b5 Merge pull request #4721 from alyssarosenzweig/ir/pool-as-you-go
Pool constants as we go
2025-07-28 11:50:08 -07:00
Ryan Houdek 6e93b0a9df Merge pull request #4725 from bylaws/sver
Windows: Force-disable SVE usage for now
2025-07-28 11:16:20 -07:00
Billy Laws a5d6d8c928 Windows: Lock invalidation tracking for the entire duration of memory ops 2025-07-28 17:46:14 +01:00
Billy Laws 561ee6601b WOW64: Wrap BTCpuSimulate to ensure correct unwinding
Some compiler versions generated FP-relative operations before loading
it from the stack, which would crash wine when APCs were used.
2025-07-28 17:24:49 +01:00
Billy Laws f44a7c95f6 Windows: Force-disable SVE usage for now 2025-07-28 17:21:57 +01:00
Tony Wasserka ac420347d1 Merge pull request #4712 from Sonicadvance1/move_legacy_binfmt
cmake: Move legacy binfmt arch-specific targets to combined
2025-07-28 09:46:10 +02:00
Tony Wasserka 493a7ccdb6 Merge pull request #4717 from Sonicadvance1/i_dislike_tuple_6
ArchHelpers: Remove pair usage in unaligned handler
2025-07-28 09:40:09 +02:00
Alyssa Rosenzweig 387db78120 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 11:42:55 -04:00
Alyssa Rosenzweig 1338a99add IR: cap constant pool
If we have more constants than registers, something will be rematerialized. Use
a simple round-robin heuristic to pick instead of the better-but-slower approach
with RA. This is a heuristic to reduce JIT time with minimal impact on code
quality. In Instcountci, the only impact is a block in oblivion only increasing
instruction count by 0.2%. And moves of constants are free for cycles at least
on Firestorm, so this isn't where we want to spend piles of JIT time anyway.

Difference at 95.0% confidence
        -0.00138911 +/- 0.00104724
        -0.418608% +/- 0.315587%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 11:42:55 -04:00
Alyssa Rosenzweig f9edd20bf6 IR: pool constants in the emitter
This is slightly worse for x87 blocks since we can't share constants between the
x87 and the main code, but otherwise should be comparable and this avoids an
expensive remapping operation.

Difference at 95.0% confidence
	-0.00474273 +/- 0.00119189
	-1.40908% +/- 0.354114%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 10:50:17 -04:00
Alyssa Rosenzweig 6b3c7319c4 IR: wrap _Constant as Constant
flag day rename/wrapping. no functional change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:33:15 -04:00
Alyssa Rosenzweig bf51fc7c36 OpcodeDispatcher: do not use Constant as an identifier
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:32:42 -04:00
Alyssa Rosenzweig e910a81c12 OpcodeDispatcher: remove sized constant use
instcountci squashed for visibility.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:26:43 -04:00
Alyssa Rosenzweig 027bd93df9 IREmitter: remove unused
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:03:00 -04:00
LC 57c0d0f25e Merge pull request #4718 from Sonicadvance1/i_dislike_tuple_7
ELFContainer: Remove tuple usage
2025-07-25 23:04:52 -04:00
Ryan Houdek 24ff6bfde5 ELFContainer: Remove tuple usage 2025-07-25 19:33:32 -07:00
LC 78770683fc Merge pull request #4716 from Sonicadvance1/i_dislike_tuple_5
FEXCore: Replace CustomIREntry tuple with struct
2025-07-25 21:38:55 -04:00
LC c0008af877 Merge pull request #4714 from Sonicadvance1/i_dislike_tuple_3
IR: Remove tuple usage from NodeIterator
2025-07-25 21:37:56 -04:00
LC bffb81241a Merge pull request #4715 from Sonicadvance1/i_dislike_tuple_4
FEXCore: Remove reference SHA implementation
2025-07-25 21:34:29 -04:00
Ryan Houdek 53ff6d54c3 Merge pull request #4720 from tobhe/readme
Readme.md: Mention Ubuntu 25.04 as supported
2025-07-25 16:40:48 -07:00
Tobias Heider 2278e3334e Readme.md: Mention Ubuntu 25.04 as supported
Support was added in 1786c2f157
2025-07-26 00:38:26 +02:00
Ryan Houdek b0c61e2b69 ArchHelpers: Remove pair usage in unaligned handler
It's being treated like optional, where a value means it has been
handled, and no value means it hasn't been handled. Stop using pair in
this case.
2025-07-25 13:01:18 -07:00
Ryan Houdek a402c308ad FEXCore: Replace CustomIREntry tuple with struct 2025-07-25 12:33:25 -07:00
Ryan Houdek 6d83a195c6 InstcountCI: Remove sha 2025-07-25 12:22:13 -07:00
Ryan Houdek 0d69e88d53 FEXCore: Remove reference SHA implementation
Due to us only enabling the CPUID extension in the case that the host
hardware supports SHA or not, this has actually been largely unused now.
Also the only hardware that doesn't support the crypto extension has
been some old Pi hardware and some other things we don't really care
about.

This code was a phenomenal reference point for implementing the SHA
versions of the instructions and would have been significantly more
difficult to implement had this not been available. Kudos to @lioncash
for having written it!

But now as we are no longer utilizing it, it is time to remove it.
2025-07-25 12:18:13 -07:00
Ryan Houdek 7309a5c3e1 HostFeatures: Fixes SHA check for vixl sim
Oops, was accidentally checking for rand.
2025-07-25 12:17:45 -07:00
Ryan Houdek 9e3c7aa9b3 IR: Remove tuple usage from NodeIterator 2025-07-25 12:10:39 -07:00
LC 30fb992cf7 Merge pull request #4708 from Sonicadvance1/geekbench_microtest
unittests/instcountci: Adds a long-lived ymm_high test
2025-07-24 21:31:04 -04:00
LC 8e5db99ba1 Merge pull request #4709 from Sonicadvance1/lock_xadd
unittests/instcountci: Adds missing lock xadd tests
2025-07-24 21:30:22 -04:00
Ryan Houdek 8103cdbec4 Merge pull request #4711 from bylaws/rwxinf
InvalidationTracker: Fix queries in the non-intersecting RWX case
2025-07-24 16:23:11 -07:00
Ryan Houdek b11835aab5 cmake: Move legacy binfmt arch-specific targets to combined
This matches the systemd path, no more _32 and _64 versions, just
binfmt_misc.

Having a mixed install where one program does one architecture and
another is weird and unsupported anyway.
2025-07-24 16:18:35 -07:00
Ryan Houdek 5b855d9baf unittests/instcountci: Adds missing lock xadd tests
Realized we were missing these, Looks like some moves could be
eliminated.
2025-07-24 15:20:02 -07:00
Billy Laws 56b6b96def InvalidationTracker: Fix queries in the non-intersecting RWX case
This needs to return the base of the non-RWX interval, with a size
that when added to the base is the start of the RWX interval.
2025-07-24 22:58:02 +01:00
Ryan Houdek 3744cad3d8 unittests/instcountci: Adds a long-lived ymm_high test
Spills in to spill-slots when it should instead spill in to the context.
2025-07-24 10:24:53 -07:00
Ryan Houdek bf8b5ed9ad Merge pull request #4670 from bylaws/callret
Implement call-ret stack optimisations
2025-07-24 10:24:30 -07:00
Billy Laws 49482fe963 InstCountCI: Update 2025-07-24 14:53:09 +01:00
Billy Laws 3497870a45 JIT: Guard lookupcache locks with the code invalidation mutex
Avoids issues with forking, as the code invalidation mutex is fork-safe.
2025-07-24 14:53:09 +01:00
Billy Laws 9a1efacefe OpcodeDispatcher: Allow for direct linking of non-multiblock direct jumps
Using an add here prevents ExitFunction from taking the direct path.

Reported by chengmingtang on Discord.
2025-07-24 14:53:09 +01:00
Billy Laws 3efb2379ee Dispatcher: Keep the call-ret stack balanced for thunk callbacks 2025-07-24 14:53:09 +01:00
Billy Laws 20efadbe66 OpcodeDispatcher: Treat ThunkOp ExitFunction as a return
ThunkOp acts as an implicit return, mark it as such so the call-ret
stack entry from the caller is popped
2025-07-24 14:53:09 +01:00
Billy Laws 31f8abfa61 Linux: Manage the call-ret stack 2025-07-24 14:53:09 +01:00
Billy Laws cf4cc71010 Dispatcher: Opportunistically perform a call-ret stack return on EC entry 2025-07-24 14:53:09 +01:00
Billy Laws 44107757a3 JIT: Rewrite block linking to support direct ExitFunction calls
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.

To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset

If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.

Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.

(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.

(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).

(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).

(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).

(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).

(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
2025-07-24 14:53:09 +01:00
Billy Laws 45ba1af388 BranchOps: Use the call-ret stack to optimise indirect ExitFunction 2025-07-24 14:53:09 +01:00
Billy Laws ba9884a26a JIT: Emit entrypoint code for call return target blocks
This is made slightly awkward by the many potential orderings of blocks
and desire to support both fallthrough jumps and calls without additional
branches.
2025-07-24 14:53:09 +01:00
Billy Laws a40d53497b OpcodeDispatcher: Emit hints for call/ret instructions 2025-07-24 14:53:09 +01:00
Billy Laws 261b7f1110 IR: Support call/ret hints in ExitFunction 2025-07-24 14:53:09 +01:00
Billy Laws c604893944 WOW64: Implement call-ret stack management 2025-07-24 14:53:09 +01:00
Billy Laws 2f16d25ab8 ARM64EC: Implement call-ret stack management 2025-07-24 14:53:09 +01:00
Billy Laws 132de3160a Windows: Implement common call-ret stack helpers
Allocate the stack with uncommited guard pages on either side, if
any faulting accesses to these occur then the call-ret SP value is
reset to the default and execution resumed.
2025-07-24 14:53:09 +01:00
Billy Laws d67b1645c9 FEXCore: Hold a frontend allocation for the call-ret stack
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
2025-07-24 14:53:09 +01:00
Billy Laws 68270ad425 LookupCache: Return whether Erase removed any cache entries 2025-07-24 14:53:09 +01:00
Billy Laws 963a8c2f08 FEXCore: Save and restore the call/ret SP from CPUState
For simplicity in cases like signal handling, always load it in
fill and store in spill, even the SP is stored in a callee save
register.
2025-07-24 14:53:09 +01:00
Billy Laws a80581e52b Arm64Emitter: Allocate a register for the call-ret SP
The host stack pointer can't be reused since explicit bounds checks
would be far too expensive, and on-stack signals prevent implicit ones
using guard pages from working.
2025-07-24 14:53:09 +01:00
Billy Laws 882cdaaeda Frontend: Move to initializing persistent sets directly 2025-07-24 14:52:54 +01:00
Billy Laws fd58f17dbe FEXCore: Track block executable ranges prior to adding cache entries
Since CodePages is now a member of the guest to host map, which could
be replaced when JITing ARM code, any additions to it must be moved after that.
Additionally there is no benefit marking code pages for invalidation at all if
they are never added to the cache as in the single-step case.

This does technically prolong the window of an existing race where guest code
modifications could be missed, however this is unlikely to cause issues and didn't
prior.
2025-07-24 14:52:54 +01:00
Billy Laws 7e5c0d7174 LookupCache: Move CodePages to GuestToHostMap
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.

The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
2025-07-24 14:52:54 +01:00
Ryan Houdek 525462e982 Merge pull request #4642 from pmatos/feature/x87-invalid-operation-bit
Implement x87 invalid operation bit on F80 mode
2025-07-23 11:05:28 -07:00
Ryan Houdek 507b3a69f0 Merge pull request #4699 from bylaws/srawow64
Inline SMC fixes and support for WOW64
2025-07-23 11:03:18 -07:00
Billy Laws b15e49910a WOW64: Support inline SMC
Matches the ARM64EC handling
2025-07-23 18:46:28 +01:00
Billy Laws 573d0858ac ARM64EC: Fix inline SMC handling for writes in the same page as the current block
Consider a page with two blocks in it, A and B. Block A performs SMC on B then A.
With the previous logic, the SMC write of B would unprotect the page and then
the inline SMC touching A would not be detected and a single-step would not be forced.
2025-07-23 18:46:28 +01:00
Billy Laws c3c2789cdf Windows: Specify the single-inst argument when calling the dispatcher 2025-07-23 18:46:28 +01:00
Billy Laws e9f14e32dd Linux: Specify the single-inst argument when calling the dispatcher 2025-07-23 18:46:26 +01:00
Billy Laws 9b65d819f4 Dispatcher: Add an argument to request a single-step on SRA fill
This already exists for ARM64EC, but this path is required for linux
and WOW64
2025-07-23 18:29:12 +01:00
Billy Laws 287344986c FEXCore: Switch ENTRY_FILL_SRA_SINGLE_INST_REG to TMP2
Will allow this to be taken as the second dispatcher argument
2025-07-23 18:29:12 +01:00
Paulo Matos 7386b4b037 instcountci: Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:16 +02:00
Paulo Matos 9a2f62c479 asm_tests: Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:16 +02:00
Paulo Matos 726656d0bf Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:10 +02:00
Ryan Houdek fe61aabc52 Merge pull request #4701 from bylaws/noexecwin
Windows: Correctly report noexec faults
2025-07-22 12:54:28 -07:00
Ryan Houdek 564e966730 Merge pull request #4698 from lioncash/float2
ASIMDOps/SVEOps: Use IsStandardFloatSize() even more
2025-07-22 12:36:16 -07:00
LC 98abfa4395 Merge pull request #4700 from bylaws/geiex
WinAPI: Avoid sign-extension of processor count in GetSystemInfo
2025-07-22 15:02:17 -04:00
Billy Laws 8e6fbe7615 Windows: Correctly report noexec faults 2025-07-22 19:30:31 +01:00
Billy Laws 27fd5505be OpcodeDispatcher: Set correct trap number for noexec faults 2025-07-22 19:30:23 +01:00
Billy Laws e383cc89eb unittests: Test for the noexec pagefault trap number 2025-07-22 19:29:17 +01:00
Alexandre Julliard a71a0ed137 WinAPI: Avoid sign-extension of processor count in GetSystemInfo 2025-07-22 19:28:36 +01:00
Lioncache ab6fc279ce ASIMDOps/SVEOps: Use IsStandardFloatSize() even more
Forgot in #4697 to take into account that there are still parts of the
emitter that qualify with its own namespace.

This also removes those where applicable to be more in line with the rest
of the cases.
2025-07-22 08:25:35 -04:00
Ryan Houdek dd5c17291a Merge pull request #4697 from lioncash/float
SVEOps: Make use of IsStandardFloatSize() more
2025-07-21 12:37:03 -07:00
Ryan Houdek 5fe06b7413 Merge pull request #4696 from lioncash/opcode
OpcodeDispatcher: Mark functions as const/static where applicable
2025-07-21 12:36:22 -07:00
Lioncache 2a8efa9891 SVEOps: Make use of IsStandardFloatSize() more
Initially introduced in #4687 as a utility for ASIMD ops, this can be used
elsewhere as well to reduce the verbosity of some other assertions.
2025-07-21 13:51:37 -04:00
LC cf43e5eaaf Merge pull request #4694 from Sonicadvance1/fix_moffset_instcountci
InstcountCI: Fix bad encoded moffset instructions
2025-07-21 10:53:26 -04:00
Lioncache cafc906b9e OpcodeDispatcher: Mark functions as const/static where applicable
Not a functional change, but makes it more obvious how these rely on object state.
2025-07-21 10:48:55 -04:00
Tony Wasserka 3387f51e24 Merge pull request #4663 from Sonicadvance1/update_format_requires
External/code-format-helper: Update requirements
2025-07-21 10:17:52 +02:00
Ryan Houdek ccf1bb26af InstcountCI: Fix bad encoded moffset instructions
These were actually encoded incorrectly, and I noticed they changed in
PR #4670 for some reason. So I dove in to the nasm source to figure out
what it takes to encode these sanely rather than with raw bytes, turns
out it's fairly trivial, just annoying to track in their parser through
iwdq->mem_offset->MEM_OFFS->qword.

Should remove the change in #4670 once it rebases.
2025-07-18 12:32:56 -07:00
Ryan Houdek fc1ca01a0b Merge pull request #4693 from neobrain/fix_libfwd_wl
LibraryForwarding/wayland: Add method signatures required by steam-runtime-launch-options
2025-07-18 10:57:50 -07:00
Tony Wasserka e508f0900c LibraryForwarding/wayland: Add method signatures required by steam-runtime-launch-options 2025-07-18 11:38:46 +02:00
LC 5123ca52ba Merge pull request #4692 from Sonicadvance1/reintroduce_cssc
FEXCore: Reintroduce support for CSSC
2025-07-17 21:21:46 -04:00
Ryan Houdek 7e1ee5bb07 FEXCore: Reintroduce support for CSSC
Now that the PF flag isn't using popcount, this is a win across the
board if the hardware supports it.

Been a while since I last looked at this, added a new instcountci file
to show the improvement.
2025-07-17 15:09:27 -07:00
LC b7ac641aa0 Merge pull request #4691 from Sonicadvance1/ignore_jemalloc_config_sources
External: Update jemallocs
2025-07-17 11:29:20 -04:00
Ryan Houdek 251ac145d9 Merge pull request #4650 from pmatos/feature/ReformatChanged
Add --changed flag to reformat.sh script
2025-07-16 23:33:29 -07:00
Ryan Houdek e17c2c8c73 Merge pull request #4661 from pmatos/feature/reformat-clang-format-19
Whole-tree reformat with clang-format-19
2025-07-16 23:33:09 -07:00
Paulo Matos 4473054d33 Add reformat sha to ignore revs for git-blame 2025-07-17 08:11:08 +02:00
Paulo Matos 5267cde60e Whole-tree reformat with clang-format-19 2025-07-17 08:10:00 +02:00
Paulo Matos 3c7ece2ef3 Update .clang-format 2025-07-17 08:09:25 +02:00
Paulo Matos 3b322d8f37 Add --changed flag to reformat.sh script 2025-07-17 08:04:41 +02:00
LC 1d4b6c6a74 Merge pull request #4689 from Sonicadvance1/disable_gcs_protection
Disable GCS in simulator and userspace
2025-07-16 20:25:16 -04:00
Billy Laws f5efe1d251 Merge pull request #4690 from Sonicadvance1/moar_padding
JIT: Add more padding
2025-07-17 00:58:37 +01:00
Ryan Houdek 09db3aba0a External: Update jemallocs
Ignore any additional jemalloc configuration options, can come from
`/etc/malloc.conf` or even the `MALLOC_CONF` environment variable.

Ensure nothing can override our options.
2025-07-16 16:35:48 -07:00
Ryan Houdek aec7deca9d Merge pull request #4688 from lioncash/hfloat2
ASIMDOps: Merge half-float 3-reg same with single/double variants
2025-07-16 15:42:08 -07:00
Ryan Houdek 8f1d4bc710 FEXLoader: Check for GCS being enabled
There is a ELF note for this but currently clang doesn't support a
`no-gcs` flag. The best we can do is check if the kernel has GCS enabled
for the current process and early exit.

Then continue to use the kernel's locking functionality to disable it if
the guest happens to try, ensuring safety.
2025-07-16 15:41:23 -07:00
Ryan Houdek 7edc8417b6 JIT: Add more padding
Go to a whole page of additional padding, #4670 adds some more size to a
block and overran the padding causing a unittest to fail.
2025-07-16 15:28:05 -07:00
Ryan Houdek 5b01682641 Arm64Emitter: Disable GCS in the simulator
FEX isn't going to be compatible with this.
PR #4670 requires this
2025-07-16 15:23:43 -07:00
Lioncache a615f8970a ASIMDOps: Constrain 3-reg same float arguments with IsQOrDRegister
These are only intended to be used with QRegisters or DRegisters, so we can constrain these
so that the proper types are enforced at compile-time.
2025-07-16 17:46:28 -04:00
Lioncache 708ea9d50a ASIMDOps: Merge half-float 3-reg same with single/double variants 2025-07-16 17:36:24 -04:00
Ryan Houdek e36a64ecc0 Merge pull request #4681 from Sonicadvance1/i_dislike_tuple_2
FEXCore/OpcodeDispatcher: Removes tuple usage for Dispatch tables
2025-07-16 14:23:18 -07:00
Ryan Houdek 82be09b6a3 FEXCore/OpcodeDispatcher: Removes tuple usage for Dispatch tables
NFC
2025-07-16 13:57:16 -07:00
Ryan Houdek eb7655d93c External/code-format-helper: Update requirements
Latest of everything, let's see what happens.
To get rid of dependabot alerts.
2025-07-16 13:54:43 -07:00
Ryan Houdek c1fe841bf3 Merge pull request #4687 from lioncash/hfloat
ASIMDOps: Merge half-float 2-reg misc with single/double variants
2025-07-16 13:50:53 -07:00
Lioncache 82ffb379c4 ASIMDOps: Merge half-float 2-reg misc with single/double variants
Unifies the interface, so that there's no need for a stark difference.

Previously, calling the single/double variant didn't require explicit template
arguments, but the half-float version did, which is inconsistent.

Technically it also made the interface more cumbersome to use in the event the
element size isn't able to be determined as a constant ahead of time.
2025-07-16 10:16:49 -04:00
Tony Wasserka befae52993 Merge pull request #4685 from pmatos/fix/clang-format-19-ignore
Ensure .clang-format-ignore is compatible with clang-format-19
2025-07-16 15:32:09 +02:00
Paulo Matos 48597d7682 Ensure .clang-format-ignore is compatible with clang-format-19
Since clang-format-19 doesn't support globstar yet, add .clang-format
to disable formatting inside External/.
2025-07-16 15:20:03 +02:00
Lioncache 05d45ccb7f Emitter: Add helper for sanitizing floating point element sizes
Will be used in a following change to reduce the amount of duplication
made in some floating point instructions.
2025-07-16 09:19:45 -04:00
LC e0ca04bfec Merge pull request #4684 from Sonicadvance1/spurious_syscall_crash
FEXCore/OpcodeDispatcher: Fix a spurious crash that can occur with multiblock
2025-07-15 21:44:51 -04:00
LC 34566861ba Merge pull request #4683 from Sonicadvance1/waitpkg_mostly_nop
FEXCore: Implement a mostly NOP implementation of waitpkg
2025-07-15 21:43:11 -04:00
Ryan Houdek dfe3b4502a FEXCore/OpcodeDispatcher: Fix a spurious crash that can occur with multiblock
If multiblock discovers a codepath with an `int 0x80` then it would
crash the emulator even if it never gets executed. Ensure that this
ERROR_AND_DIE_FMT instead just gets handled as an UnhandledOp to ensure
the Core early terminates the block.

Found by having steamwebhelper spuriously crash when it hit this.
Adds a simple unittest to ensure discovery doesn't break again.
2025-07-15 17:43:16 -07:00
Ryan Houdek 4884aeef20 Config: Adds option to disable wfxt 2025-07-15 16:38:15 -07:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Ryan Houdek 29fccd0e1b Merge pull request #4682 from bylaws/asahi
Windows: Support enabling hardware TSO on Asahi Linux
2025-07-15 14:25:03 -07:00
Ryan Houdek 35e4ac5a64 Merge pull request #4668 from Sonicadvance1/implement_nx
FEXCore: Implement support for NX bit.
2025-07-15 12:40:48 -07:00
Ryan Houdek 43bba77840 FEXCore: Implement support for NX bit.
Long time coming but thanks to bylaw's changes in #4474, this is now
trivial to implement.

Fixes #2175
2025-07-15 10:41:24 -07:00
Billy Laws 0361274366 Windows: Support enabling hardware TSO on Asahi
Relies on a wine-side patch to expose the prctl to the PE-side
2025-07-15 17:45:21 +01:00
Ryan Houdek 5b0e703811 Merge pull request #4678 from neobrain/refactor_unuse_unused
Drop unnecessary uses of maybe_unused
2025-07-15 09:03:55 -07:00
Ryan Houdek 05742a8e28 Merge pull request #4677 from neobrain/refactor_misc_logs
Miscellaneous log changes
2025-07-15 09:03:12 -07:00
Ryan Houdek 170fbb596c Merge pull request #4679 from neobrain/fix_signed_overflow
Arm64Emitter: Fix signed overflow
2025-07-15 09:02:29 -07:00
Ryan Houdek b72bdf9948 Merge pull request #4680 from lioncash/vfaddv
VectorOps: Correct benign VFAddV op cast
2025-07-15 09:00:41 -07:00
Lioncache 678e03b470 VectorOps: Correct benign VFAddV op cast
This was using the non-float VAddV variant, but had the same behavior,
since the fields were named the same. So this is just a correctness fix.
2025-07-15 11:41:28 -04:00
Tony Wasserka 0f45f3a24d Arm64Emitter: Fix signed integer overflows 2025-07-15 17:17:48 +02:00
Tony Wasserka 9b503f3702 IRDumper: Remove unnecessary use of maybe_unused 2025-07-15 17:10:37 +02:00
Tony Wasserka 7f3f619518 GDBJIT: Remove unnecessary use of maybe_unused 2025-07-15 17:10:27 +02:00
Tony Wasserka f7d59880cc ELFCodeLoader: Drop unused maybe_unused attribute 2025-07-15 17:07:36 +02:00
Tony Wasserka ec9e266c3c FDUtils: Mark get_fdpath as nodiscard 2025-07-15 17:07:36 +02:00
Tony Wasserka ed62c02494 AllocatorOverride: Make error message more prominent 2025-07-15 16:55:18 +02:00
Tony Wasserka d5b166bdd9 OpcodeDispatcher: Strengthen error message to fatal 2025-07-15 16:55:18 +02:00
Tony Wasserka 648c726bb1 LogManager: Use assert log level for ERROR_AND_DIE_FMT 2025-07-15 16:55:18 +02:00
Tony Wasserka 88f30afb26 Merge pull request #4672 from neobrain/refactor_format_formatter
code-format-helper: Fix formatting for the format helper
2025-07-15 14:27:22 +02:00
LC e6f319d6b3 Merge pull request #4674 from neobrain/refactor_todrop
Drop TODO defines
2025-07-15 07:12:35 -04:00
LC a80d8b15c2 Merge pull request #4673 from neobrain/refactor_value_nor
XXFileHash: Drop unnecessary use of value_or
2025-07-15 07:09:47 -04:00
LC 2a372895e4 Merge pull request #4671 from Sonicadvance1/xsaveopt
FEXCore: Implement xsaveopt
2025-07-15 07:08:15 -04:00
Tony Wasserka 3b67e573b5 Revert "FexHeaderUtils: Add TodoDefines"
This reverts commit ad1fd7f54b.
2025-07-15 10:04:01 +02:00
Tony Wasserka a3a55d19b8 Revert "FEX_TODO: Convert some XXX to FEX_TODO"
This reverts commit 256df76674.
2025-07-15 10:03:56 +02:00
Tony Wasserka 4e25bce616 FEXServer: Drop use of FEX_TODO macro 2025-07-15 10:03:49 +02:00
Tony Wasserka 135477e539 XXFileHash: Drop unnecessary use of value_or 2025-07-15 09:47:52 +02:00
Tony Wasserka 147b1f2293 code-format-helper: Replace spurious tab with spaces 2025-07-15 09:43:55 +02:00
Ryan Houdek 401dc69586 InstcountCI: Add xsaveopt
Just in-case it changes.
2025-07-14 17:32:45 -07:00
Ryan Houdek 0dff992fab FEXCore: Implement xsaveopt
It's the same as xsave because we don't track a hidden "xinuse" hardware
mask. So this is a trivial implementation.
2025-07-14 17:31:47 -07:00
LC 4f605739f1 Merge pull request #4667 from Sonicadvance1/i_dislike_tuple
XXFileHash: Remove a tuple usage
2025-07-14 18:21:54 -04:00
LC e71e10adcb Merge pull request #4669 from Sonicadvance1/handful_cpuinfo_missing
EmulatedFiles/cpuinfo: Add a few missing flags
2025-07-14 18:21:03 -04:00
Ryan Houdek 9f1bc05644 EmulatedFiles/cpuinfo: Add a few missing flags
bus_lock_detect might be interesting to expose if applications ever
change behaviour depending on the underlying fault behaviour.
2025-07-14 14:56:24 -07:00
Ryan Houdek 603878b3e9 XXFileHash: Remove a tuple usage 2025-07-14 13:09:27 -07:00
Ryan Houdek 70ce323537 Merge pull request #4666 from pmatos/fix/git-clang-format-19
Set clang-format-19 as the version git-clang-format should run in wor…
2025-07-14 10:38:14 -07:00
Paulo Matos b5f8dd93f7 Set clang-format-19 as the version git-clang-format
Unfortunately git-clang-format-19 will not call clang-format-19 but the system clang-format
so we need to hardcode it here.
2025-07-14 18:37:47 +02:00
Ryan Houdek 2e7b86e452 Merge pull request #4474 from bylaws/mentry
Support multiple entrypoints into a multiblock and executable permission tracking
2025-07-11 19:05:42 -07:00
Ryan Houdek b2eeaf75a9 Merge pull request #4659 from bylaws/sysfix
ARM64EC: Rely on syscall export sorting for deriving their IDs
2025-07-11 18:09:07 -07:00
Ryan Houdek 5d4bddcfb9 Merge pull request #4662 from pmatos/patch-2 2025-07-11 08:49:45 -07:00
Paulo Matos 13ef676115 Do not reformat files in External/ 2025-07-11 15:54:02 +02:00
Ryan Houdek 67bfd3877c Merge pull request #4651 from pmatos/feature/ClangFormat19Upgrade
Upgrade to clang-format-19
2025-07-11 01:03:49 -07:00
Ryan Houdek 2be619b011 Merge pull request #4658 from bylaws/fmt
External: Update libfmt to master to support clang 21
2025-07-10 11:08:33 -07:00
Ryan Houdek ebd0a3ceaf Merge pull request #4660 from lioncash/null
VixlUtils: Fix null pointer dereference vector in IsImmLogical()
2025-07-10 11:07:01 -07:00
Lioncache d8da14d550 VixlUtils: Fix null pointer dereference vector in IsImmLogical()
Technically we can end up doing null pointer dereferencing here since checks were being
chained with || instead of &&. So we can just separate the checks out.
2025-07-10 12:09:20 -04:00
Billy Laws 076c156cbb Frontend: Warn on invalid instructions in entry blocks 2025-07-10 16:51:34 +01:00
Billy Laws 6eaaab8dc2 IntervalList: Also return the full matching interval on query 2025-07-10 16:51:34 +01:00
Billy Laws 46bc8e499a Frontend: Always treat FEXCore X86 callbacks as executable 2025-07-10 16:51:34 +01:00
Billy Laws 21865d09a2 SyscallsSMCTracking: Handle READ_IMPLIES_EXEC for executable mapping queries 2025-07-10 16:51:34 +01:00
Billy Laws 31903d0c0b Frontend: Treat instructions in non-executable memory as invalid 2025-07-10 16:51:33 +01:00
Paulo Matos c7b7cbcac1 Upgrade to clang-format-19
Fixes #4577
2025-07-10 17:11:44 +02:00
Billy Laws d0af858a9c LinuxEmulation: Implement SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 2e5c6283f5 CodeSizeValidation: Add dummy QueryGuestExecutableRange impl 2025-07-10 16:00:24 +01:00
Billy Laws 261df26cff DummyHandlers: Stub SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 96bd6ce3e5 Windows: Implement SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 14eed89bb6 SyscallHandler: Add method to query executable memory ranges 2025-07-10 16:00:24 +01:00
Billy Laws 3feb354186 Frontend: Keep the associated thread object as a member
Avoids an additional layer of indirection for callbacks. Passing them
around deep into instruction decoding logic doesn't provide much benefit
seeing as there will always be one frontend object per thread.
2025-07-10 16:00:24 +01:00
Billy Laws d41cb3b69d FEXCore: Support multiple entrypoints into a multiblock
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
2025-07-10 16:00:24 +01:00
Billy Laws cdef1ed0c5 Frontend: Explore after call instructions with multiblock
Each instruction after a call instruction can be treated as an
additional entrypoint to the multiblock
2025-07-10 16:00:24 +01:00
Billy Laws 4f79acc32a X86Tables: Mark call instructions with a flag 2025-07-10 16:00:24 +01:00
Billy Laws 9328b099c3 CMake: Set CMAKE_AR for MinGW toolchains
Avoids the need for the toolchain to override the system AR.
2025-07-10 16:00:24 +01:00
Billy Laws 378d0351cf External: Update libfmt to master to support clang 21 2025-07-10 16:00:24 +01:00
Billy Laws 793e3adbb4 ARM64EC: Rely on syscall export sorting for deriving their IDs
Rather than relying on wine-specific alphabetical behaviour, that
has since been changed. Rely on their addresses being sorted which
is more stable behaviour also present in Windows.
2025-07-10 16:00:07 +01:00
Billy Laws 2f4f4253e7 External: Update libfmt to master to support clang 21 2025-07-10 15:52:17 +01:00
Ryan Houdek 10c69d561e Merge pull request #4656 from bylaws/mingw-new
CMake: Set CMAKE_AR for MinGW toolchains
2025-07-09 19:26:30 -07:00
Billy Laws bbf6e8e0c5 CMake: Set CMAKE_AR for MinGW toolchains
Avoids the need for the toolchain to override the system AR.
2025-07-10 00:33:37 +01:00
Ryan Houdek d8a4d03501 Merge pull request #4655 from alyssarosenzweig/bug/divisor-mask
Fix divisor masking
2025-07-09 12:15:16 -07:00
Alyssa Rosenzweig 8ecc8ef75d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Alyssa Rosenzweig cfd4318080 unittests: add 32-bit masking divide unit test
fails on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Alyssa Rosenzweig d23bb01e96 JIT: fix divisor masking
oversight. should fix Steam.

Fixes: de4becc26 ("OpcodeDispatcher: mask certain divisors")
Closes: #4652
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Ryan Houdek 60f63ca467 Merge pull request #4649 from bylaws/rexvex
Frontend: Raise #UD on invalid VEX/REX encodings
2025-07-09 10:37:50 -07:00
LC 40378cd4d8 Merge pull request #4654 from Sonicadvance1/vcvtsd2si_test
unittests: Update vcvtsd2si test
2025-07-09 12:16:02 -04:00
Ryan Houdek f44af48751 Merge pull request #4648 from bylaws/sfmodreg
SecondaryGroupTables: Specify FLAGS_SF_MOD_REG_ONLY for more ops
2025-07-09 09:02:12 -07:00
Ryan Houdek d01a48b773 unittests: Update vcvtsd2si test
This would have failed prior to #4647 getting merged.
2025-07-09 08:50:19 -07:00
Ryan Houdek bd136eae71 Merge pull request #4647 from bylaws/sizrfix
OpcodeDispatcher: Fix several cases where incorrectly sized loads/stores could be used
2025-07-09 08:49:38 -07:00
LC e669370629 Merge pull request #4646 from bylaws/uaffix
PoolBufferWithTimedRetirement: Unclaim in dtor
2025-07-09 10:14:13 -04:00
Billy Laws 26215d5425 VEXTables: Complete decode flag information 2025-07-09 00:53:23 +01:00
Billy Laws baa5c6f79b Frontend: Require VEX.vvvv is 0 when unused 2025-07-09 00:53:23 +01:00
Billy Laws c0fcb27aa5 Frontend: Add instruction flags to specify valid REX.W encodings 2025-07-09 00:53:23 +01:00
Billy Laws 02b7938b93 Frontend: Add instruction flags to specify valid VEX.L encodings 2025-07-09 00:53:23 +01:00
Billy Laws 00fec3d51f Update InstCountCI 2025-07-09 00:52:59 +01:00
Billy Laws 0f6ceaac05 OpcodeDispatcher: Load only the element size from memory for VFMAScalarImpl 2025-07-09 00:52:59 +01:00
Billy Laws 635816c07f OpcodeDispatcher: Force ElementSize loads for UCOMISxOp 2025-07-09 00:52:59 +01:00
Billy Laws 2fac5c23ff OpcodeDispatcher: Always use 32-bit load/store for {LD,ST}MXCSR 2025-07-09 00:52:59 +01:00
Billy Laws ff4c1cf0d5 X86Tables: Fix (V)CV(T)TSD2SI size flags
This led to incorrect OOB handling for 32-bit dests.
2025-07-09 00:52:45 +01:00
Billy Laws c2e2d1e92f OpcodeDispatcher: Only read at most ElementSize in CVTFPR_To_GPR 2025-07-09 00:50:47 +01:00
Billy Laws 0ab56bad95 SecondaryGroupTables: Specify FLAGS_SF_MOD_REG_ONLY for more ops 2025-07-08 23:45:32 +01:00
Billy Laws 407c5a0f78 PoolBufferWithTimedRetirement: Unclaim in dtor
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.

Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
2025-07-08 23:38:43 +01:00
Ryan Houdek 3ba84ad06a Docs: Update for release FEX-2507 2025-07-07 23:49:56 -07:00
Ryan Houdek c6aae9e05a Merge pull request #4634 from Sonicadvance1/fix_horizon
EmulatedFiles: Emulate `current_clocksource`
2025-07-07 21:16:54 -07:00
Ryan Houdek 95b4618833 Merge pull request #4644 from ChanthMiao/fix/sigframe_mistake
Fix: wrong magic value in fpstate.
2025-07-06 17:07:31 -07:00
Changwei Miao 1fe17d55d9 Fix: wrong magic value in fpstate.
FEX should only set fpx_sw_bytes.magic1 with FP_XSTATE_MAGIC when
avx is enabled. Otherwise it may cause segfault in ntdll::save_context,
which requires access to extended xstate info if magic1 equals FP_XSTATE_MAGIC.

Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2025-07-06 18:44:37 +08:00
Ryan Houdek ead73371d9 EmulatedFiles: Emulate current_clocksource
WINE uses this to determine TSC frequency and because it doesn't say
`tsc` on ARM devices, it was ignoring TSC and instead using CPU maximum
frequency.

This was causing Horizon to think the TSC ran at whatever the max
frequency of a core was  (1.8Ghz to 2.6Ghz depending?) This was causing
all of Horizon Zero Dawn's physics to run at slower than real time
speeds because our 1Ghz (on Orion) TSC is significantly lower than the
max clock speeds of the cores.

This is still a bug in Wine that it is using the maximum CPU clock speed
in the case of current_clocksource not being TSC, but that's a battle
for a different time.
2025-07-05 21:10:53 -07:00
Ryan Houdek 6f089a4323 Merge pull request #4643 from tstellar/llvm-21
Fix build with LLVM >= 21
2025-07-05 16:51:38 -07:00
Ryan Houdek 640f024551 Merge pull request #4641 from Sonicadvance1/optimize_sincos
JIT: Optimize x87 FSINCOS
2025-07-05 16:26:37 -07:00
Tom Stellard 99920f89dd Fix build with LLVM >= 21 2025-07-05 16:50:34 +00:00
Ryan Houdek 8ea276267f InstcountCI: Update 2025-07-03 18:05:52 -07:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Ryan Houdek fa0a54deb9 Merge pull request #4640 from alyssarosenzweig/bug/fix-hades
Fix Hades
2025-07-03 14:53:01 -07:00
Alyssa Rosenzweig 046043090f unittests: add move merging test
this hits a nasty case with post-RA merging. fails on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:56 -04:00
Alyssa Rosenzweig 61150a18cc RegisterAllocationPass: fix bookkeeping with merging
this fixes Hades.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:53 -04:00
Ryan Houdek 95ca20cfee Merge pull request #4636 from alyssarosenzweig/bug/ra-invariant
RegisterAllocationPass: assert an invariant in post-RA prop
2025-07-02 18:17:08 -07:00
Alyssa Rosenzweig 5a536d47fd RegisterAllocationPass: assert an invariant in post-RA prop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:53:47 -04:00
Ryan Houdek afbc7da027 Merge pull request #4629 from alyssarosenzweig/opt/cpuid-basic
Optimize some constant cpuid/xgetbv cases
2025-07-02 10:25:01 -07:00
Alyssa Rosenzweig c093c08c40 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 360d8c629e RegisterAllocationPass: optimize cpuid
for constant function where we don't have a leaf. this isn't fully general but
we can't do better without a more general post-RA optimizer. i'm not inclined to
do that unless/until we get hot blocks demonstrating its value (that we can
compare against the JIT time hit of the heavier-duty optimizer.)

however this special case we can (and should) optimize for now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig abb41d39e4 RegisterAllocationPass: optimize xgetbv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig d4eb4ef594 IR: plumb CPUID into RA pass
for cpuid folding.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 02f45854e8 IR: include a fence in CPUID
easier for post-RA to chew thru.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Ryan Houdek 62de1004df Merge pull request #4635 from neobrain/fix_libfwd_wl_more
LibraryForwarding/wayland: Add new interface objects
2025-07-02 09:01:11 -07:00
Tony Wasserka 7f216ca02f LibraryForwarding/wayland: Add new interface objects 2025-07-02 16:51:03 +02:00
LC 492b0fdda8 Merge pull request #4632 from neobrain/feature_nix
Build: Add nix-based helpers to facilitate cross-compilation
2025-07-01 16:23:58 -04:00
LC bb072c0112 Merge pull request #4633 from Sonicadvance1/noexec_test
unittests: Adds unittest for no-exec testing
2025-07-01 16:20:44 -04:00
Ryan Houdek afabe7cb47 Merge pull request #4627 from alyssarosenzweig/opt/long-div-peephole-ready
Optimize long division
2025-06-30 13:59:48 -07:00
Ryan Houdek 38e0fc2434 unittests: Adds unittest for no-exec testing
In preparation for #4474
2025-06-30 13:19:33 -07:00
Alyssa Rosenzweig 16a70eafc6 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:59:27 -04:00
Alyssa Rosenzweig de4becc26e OpcodeDispatcher: mask certain divisors
needed for fusing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig af23f4325f OpcodeDispatcher: reorder xor-with-self sequence
this lets us peephole fuse things even when there are flags calculated in the
way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Ryan Houdek 94af96df8f Merge pull request #4631 from neobrain/fix_libfwd_wl_cutter
LibraryForwarding: Fix various Wayland issues
2025-06-30 11:38:13 -07:00
Ryan Houdek e685ab818e Merge pull request #4619 from neobrain/refactor_drop_config_h_in
Remove code generation build step for install prefix
2025-06-30 11:35:26 -07:00
Ryan Houdek 5f2a72b65b Merge pull request #4624 from Sonicadvance1/static_analysis_fixes
Some static code analysis fixes
2025-06-30 11:28:06 -07:00
Tony Wasserka c1842a6167 Build: Add nix-based helpers to manage toolchains for ARM64EC/WOW64
The cmake_configure_woa*.sh scripts will automatically install any required
cross-toolchains required to enable either ARM64EC or WOW64 builds of FEX,
and it will initialize the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell WineOnArm/shell.nix`,
which will make the toolchain available via environment variables. This also
generates a meson crossfile for building VKD3D or vkd3d-proton.
2025-06-30 16:50:13 +02:00
Tony Wasserka 3f3907b5d1 Build: Add nix-based helpers to manage toolchains for FEXLinuxTests
The cmake_enable_flt.sh script will automatically install the required
cross-toolchains required to build FEXLinuxTests, and it will reconfigure
the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell FEXLinuxTests/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 1899465390 Build: Add nix-based helpers to manage toolchains for library forwarding
The cmake_enable_libfwd.sh script will automatically install any required
cross- toolchains and development headers required to enable library
forwarding, and it will reconfigure the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell LibraryForwarding/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 18360d4ccb LibraryForwarding/wayland: Add more method signatures 2025-06-27 10:57:10 +02:00
Tony Wasserka a691c3cd99 LibraryForwarding/wayland: Fix mprotect call when crossing page boundaries 2025-06-27 10:57:10 +02:00
Tony Wasserka 70bc561bbf LibraryForwarding/wayland: Fix wl_proxy_marshal_array_constructor 2025-06-27 10:57:10 +02:00
Tony Wasserka b9222d8431 LibraryForwarding/unittests: Fix build with clang 20 2025-06-27 10:57:10 +02:00
Tony Wasserka a5fad89e57 Merge pull request #4621 from neobrain/feature_update_readme
Update Readme.md
2025-06-26 21:54:53 +02:00
Tony Wasserka a14360b89d Update Readme.md 2025-06-26 21:41:23 +02:00
Tony Wasserka c9aaedd217 CPack: Update package description 2025-06-26 21:41:23 +02:00
Alyssa Rosenzweig cda15ce9ea RegisterAllocationPass: optimize long divsion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 62410c4381 InstCountCI: add another udiv case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:22 -04:00
Ryan Houdek cf82b56dd8 CPUBackend: Remove unused variable
CID 482006
2025-06-19 16:51:30 -07:00
Ryan Houdek 4a74bea7ab InstcountCI: Don't use a global static initializer for CodeSizeValidation
Relies on fmt facet initialization order which isn't guaranteed to have
correct initialization order.

CID 482003
2025-06-19 16:51:26 -07:00
Ryan Houdek c6d8e60ef8 OpcodeDispatcher: Make sure to initialize ArithRef
CID 482002
2025-06-19 16:46:15 -07:00
Ryan Houdek 1212cd526a VDSOEmulation: Sanitize sysconf result
Unlikely to fail but make sure.

CID 482021
2025-06-19 16:43:53 -07:00
Ryan Houdek 79a8ed53b6 CPUBackend: Make sure to zero initialize variable
CID 482022
2025-06-19 16:41:24 -07:00
Ryan Houdek a6ce115d9c FEXLoader: Handle bad LogFile path
Just go silent but print a log about it.

CID 482031
CID 482023
2025-06-19 16:40:03 -07:00
Ryan Houdek 646a5a7f9e CodeEmitter/ASIMD: Removes redundant ternary selection
Redundant and unnecessary.

CID 482017
2025-06-19 16:31:08 -07:00
Ryan Houdek 7f71b6f1b2 Passes: Use move instead of copy semantics
To initialize in-place.

CID 482035
2025-06-19 16:22:39 -07:00
Ryan Houdek 16d5ca447f FEXCore: DebugData is never null now
This is a required data structure to exist.

CID 482036
2025-06-19 16:21:19 -07:00
Ryan Houdek 3d0c20a263 Merge pull request #4615 from neobrain/feature_logging_qol
Improve log message formatting
2025-06-19 12:50:44 -07:00
Ryan Houdek 1c38b8b046 Merge pull request #4614 from alyssarosenzweig/opt/easy-mov-elim
Merge moves that are immediately consumed
2025-06-19 12:26:07 -07:00
Alyssa Rosenzweig e1124480be InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 5243f50ed1 RegisterAllocationPass: merge 32-bit mov + 64-bit and
mov wA, wB
  and xA, xA, ...

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig cde805147f RegisterAllocationPass: merge 32-bit moves
mov wA, wB
  op wA, wA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 7e39eb3df2 RegisterAllocationPass: merge full size moves
mov xA, xB
  op xA, xA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig af366d4480 RegisterAllocationPass: skip inlineconstant in RA
similar reasoning as guestopcode.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 33ef98aae7 RegisterAllocationPass: refactor push/pop merge
to make way for move merging.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig c16db2db4a RegisterAllocationPass: simplify an expression
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 059d980c33 IR: give StoreRegister a precoloured destination
this will eliminate an annoying special case in post-RA opts.

No difference proven at 95.0% confidence

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 581381fd86 IR: make 0 the invalid physical register
so zero init works as expected

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Ryan Houdek 0072b289bb Merge pull request #4622 from neobrain/feature_fexconfig_logo
FEXConfig: Add icon
2025-06-18 01:46:30 -07:00
Tony Wasserka e2bd79087e FEXConfig: Add icon 2025-06-18 10:33:07 +02:00
Ryan Houdek 6cd78fc90d Merge pull request #4618 from neobrain/feature_infer_32bit_libfwd_paths
LibraryForwarding: Infer folders for 32-bit wrappers automatically
2025-06-17 17:55:38 -07:00
Ryan Houdek c346aca241 Merge pull request #4620 from neobrain/refactor_data_files
Move data files to Data/
2025-06-17 17:54:28 -07:00
Ryan Houdek bf83569f0b Merge pull request #4623 from neobrain/feature_tracy_0_12
External: Update Tracy submodule to version 0.12.1
2025-06-17 17:53:47 -07:00
Tony Wasserka 1502f04a8a External: Update Tracy submodule to version 0.12.1
Notably, this update adds flame graph functionality for aggregated data.
2025-06-17 17:51:52 +02:00
Tony Wasserka 7bd9d0ae23 Data: Move Dockerfile 2025-06-17 16:40:42 +02:00
Tony Wasserka cf57afdf26 Data: Move CI folder to Data/CI 2025-06-17 16:40:42 +02:00
Tony Wasserka 43e6aebc7a Data: Move CMake support scripts to Data/CMake 2025-06-17 16:40:42 +02:00
Tony Wasserka 22780993e1 Data: Move CPack files to Data/CMake/ 2025-06-17 16:40:42 +02:00
Tony Wasserka 578dcee9af Data: Move toolchain files to a central location 2025-06-17 16:40:42 +02:00
Tony Wasserka 61d77e3f9b LibraryForwarding: Infer folders for 32-bit wrappers automatically
There's no need to bother the user to select these paths manually.
Instead, just use the same folder names with _32 appended.

Fixes #4588.
2025-06-17 11:57:35 +02:00
Tony Wasserka 4ce0acba80 Remove now unused Config.h.in 2025-06-17 11:40:32 +02:00
Tony Wasserka d137212222 FEXGetConfig: Infer install prefix from executable path 2025-06-17 11:40:32 +02:00
Tony Wasserka 9ad4e3a6a0 FEXBash: Clean up and fix FEXInterpreter lookup
Previously, the first attempt to look up a FEXInterpreter would always
fail due to a missing path separator.

Additionally, fallback lookup now uses /proc/self/exe to find a path relative
to the FEXBash executable. The previous use of FindContainerPrefix does not
seem to be required anymore in current Steam versions.
2025-06-17 11:40:32 +02:00
Tony Wasserka 57627d4fcf LogManager: Drop unused STDOUT/STDERR log levels 2025-06-16 13:54:03 +02:00
Tony Wasserka 23b69271eb Use consistent log message formatting for all modules 2025-06-16 13:54:03 +02:00
LC 9d2f557666 Merge pull request #4616 from alyssarosenzweig/ici/32bit-div
InstructionCountCI: add 32-bit division cases
2025-06-13 13:18:55 -04:00
Alyssa Rosenzweig 4f9e352ff0 InstructionCountCI: add 32-bit division cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-13 13:05:35 -04:00
Tony Wasserka 3e85e60a30 LogManager: Use colors for logging when possible 2025-06-13 15:14:47 +02:00
Tony Wasserka febce21b21 FEXServer: Separate PID and TID with a pipe instead of a period
Since most software considers the former a word boundary but not the latter,
this allows the individual values to be copy-pasted more conveniently.
2025-06-13 14:54:22 +02:00
Tony Wasserka 755364e2df FEXServer: Clean up time display for logging
This is now relative to the time of the first message. Furthermore, display
precision is limited to milliseconds (which are actually zero-padded now!).
2025-06-13 14:54:13 +02:00
Tony Wasserka df63979773 LogManager: Avoid ^C being printed when quitting foreground FEXServer 2025-06-13 14:31:25 +02:00
Tony Wasserka 992d86bbc1 LogManager: Shorten debug level strings to a single letter 2025-06-13 14:31:25 +02:00
Ryan Houdek 534b338161 Merge pull request #4612 from alyssarosenzweig/ir/merge-divrem-2
IR: merge integer division & remainder
2025-06-12 14:29:02 -07:00
Alyssa Rosenzweig ef250f936c InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig e57130e364 JIT: drop UDiv extensions
we already extend in the dispatcher.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig eedcb35270 IR: merge ldiv/lrem handlers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig 1eb470083c IR: merge div/rem opcodes
it's simpler & faster to calculate both together, matching the x86 semantic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-10 17:08:20 -04:00
LC 1f15a4e35b Merge pull request #4611 from Sonicadvance1/ubisoft_ptrace
Linux: Implement enough of ptrace to allow Ubisoft launcher
2025-06-10 11:35:49 -04:00
Ryan Houdek 7ed9bea16b Linux: Implement enough of ptrace to allow Ubisoft launcher
Wine does some minimal attaching, peeking, poking, and detaching to have
a different process inspect another's TEB region. Ubisoft's launcher
does this to check if a debugger is attached and rejects it in the case
that it is.

While this implementation isn't all-encompassing, it is good enough for
this family of games.
2025-06-09 16:32:42 -07:00
Ryan Houdek 8b1383d235 Docs: Update for release FEX-2506 2025-06-04 10:48:24 -07:00
LC a73fab3bb5 Merge pull request #4609 from Sonicadvance1/add_plucky_install
Scripts/InstallFEX: Add plucky
2025-06-03 16:30:45 -04:00
Ryan Houdek 7bd64d9c53 Merge pull request #4608 from alyssarosenzweig/ici/case
InstructionCountCI: fix sdiv case
2025-06-03 12:50:36 -07:00
Ryan Houdek 1786c2f157 Scripts/InstallFEX: Add plucky
This was added to the PPA last month.
2025-06-03 12:31:40 -07:00
Alyssa Rosenzweig 958b671736 InstructionCountCI: fix sdiv case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:30:42 -04:00
Ryan Houdek 6549b66cf6 Merge pull request #4606 from alyssarosenzweig/opt/cdq
Optimize CDQ
2025-06-03 12:24:15 -07:00
Alyssa Rosenzweig 0c855a5ce3 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:39 -04:00
Alyssa Rosenzweig 6864d48dcf OpcodeDispatcher: optimize cdq
prereq to optimizing sign-ext+ldiv in a reasonable way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:06 -04:00
Ryan Houdek 99816a23a8 Merge pull request #4583 from neobrain/refactor_syscalls_unify
LinuxSyscalls: Reduce code duplication between 32-bit and 64-bit paths
2025-06-03 09:54:09 -07:00
Ryan Houdek 2ff9546523 Merge pull request #4599 from pmatos/FEXServerFind
Fixes FEXServer path search
2025-06-03 09:52:39 -07:00
Ryan Houdek 06541f21d6 Merge pull request #4605 from neobrain/fix_cmake_full_libdir2
Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values
2025-06-03 09:52:26 -07:00
Tony Wasserka 45a37edd4a Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values 2025-06-03 17:45:29 +02:00
Ryan Houdek 2a713a1f51 Merge pull request #4604 from neobrain/fix_cmake_install_prefix
CMake: Generate DATA_DIRECTORY dynamically unless explicitly set
2025-06-03 08:43:24 -07:00
Tony Wasserka 3e104de377 CMake: Generate DATA_DIRECTORY dynamically unless explicitly set
This allows changes to CMAKE_INSTALL_PREFIX to be automatically picked up and
propagated properly. Previously, you had to update 4 variables in 3 files by
hand to do so.
2025-06-03 16:58:06 +02:00
Paulo Matos b22b316e70 Removes handling of FEXServer from the FEXBash wrapper
This was not necessary since FEXInterpreter is already doing it,
and doing it properly. The previous implementation if FEXBash was
incomplete.
2025-06-03 12:24:03 +02:00
Tony Wasserka df461546c5 LinuxSyscalls: Make error return values consistent 2025-06-03 11:10:08 +02:00
Tony Wasserka 1c1c43cd86 LinuxSyscalls: Fix formatting 2025-06-03 11:10:08 +02:00
Tony Wasserka 3eac9f937e LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:08 +02:00
Tony Wasserka f13a0d8e84 LinuxSyscalls: Unify shmdt implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka ca12dc9213 LinuxSyscalls: Unify shmat implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 365ed2cd70 LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 7efbfed0bf LinuxSyscalls: Unify mmap and munmap implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka ed502738c6 LinuxSyscalls: Make Get32BitAllocator interface virtual 2025-06-03 11:10:07 +02:00
Ryan Houdek ba162bb058 Merge pull request #4603 from neobrain/fix_cmake_full_libdir
LibraryForwarding: Use CMAKE_INSTALL_FULL_LIBDIR instead of constructing the library install paths manually
2025-06-02 15:49:22 -07:00
Tony Wasserka f194d35913 LibraryForwarding: Use CMAKE_INSTALL_FULL_LIBDIR instead of constructing the library install paths manually
This fixes issues in the nix build, where CMAKE_INSTALL_LIBDIR is an
absolute path instead of the typical "lib(64)".
2025-06-03 00:37:23 +02:00
Ryan Houdek e81c84e00a Merge pull request #4602 from Sonicadvance1/add_instcountci_tests
InstcountCI: Adds tests for instructions discovered by #4597
2025-06-02 12:45:51 -07:00
Ryan Houdek 68dc9030bc Merge pull request #4601 from Sonicadvance1/fix_futimesat_flake
unittests/futimesat: Fixes flake
2025-06-02 12:45:26 -07:00
Ryan Houdek fedad275e7 Merge pull request #4597 from alyssarosenzweig/opt/drop-pile-of-constprop
Constant fold on the fly
2025-06-02 12:34:53 -07:00
Ryan Houdek f6b4c76d76 InstcountCI: Adds tests for instructions discovered by #4597
Apparently I completely missed that cpuid, xgetbv, syscall,
l{u,}{div,rem} were failing to hit their optimized cases for inlining
and avoiding 128-bit software divide.

The divisions are a clear performance regression for 64-bit applications
since that is the only real way to do a 64-bit division on x86, I added
those specifically because it sped up games.

CPUID depends heavily on the game, since some games use that as a
serialization instruction fairly heavily.

XGETBV is trivial since it matches behaviour of CPUID (and is basically
an extension of it).

Syscall inlining can save a decent amount of time, again heavily depends
on game.

Adds multi-inst tests for all of these situations so that once it gets
fixed (Apparently broken once RCLSE got stripped out), we can see that
they keep working. Obviously tests couldn't have existed in instcountCI
before since we didn't support multi-instruction tests.
2025-06-02 12:10:39 -07:00
Alyssa Rosenzweig 27854aa091 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig d966ae145e ConstProp: merge inline + pooling
now that the algebraic/folding opts are gone, we can do this in one pass for a
2.5% speedup:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44474704    0.46750433    0.45455258    0.45446569  0.0044727894
+  50    0.43149892    0.45173984    0.44252267    0.44295575  0.0045621814
Difference at 95.0% confidence
	-0.0115099 +/- 0.00179263
	-2.53263% +/- 0.394447%
	(Student's t, pooled s = 0.00451771)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig 4cb37e6a1b ConstProp: drop constant folding and algebraic opts
No longer needed.

The total difference from the beginning of this series (all the prep work to
make this change possible) plus this commit is a modest 0.4% win.

    N           Min           Max        Median           Avg        Stddev
x 100     0.4472467    0.46646308    0.45708424    0.45713057  0.0040838243
+ 100    0.44707586    0.46581227    0.45479448    0.45509309  0.0037548573
Difference at 95.0% confidence
	-0.00203748 +/- 0.00108734
	-0.445711% +/- 0.237862%
	(Student's t, pooled s = 0.00392279)

...in addition to a net deletion of 144 lines of code.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:24:58 -04:00
Alyssa Rosenzweig d9da81e99b ConstProp: drop dead cross-instr opts
these can't work due to architectural limitations. they could be ported to
post-RA passes, I think, but having them here now is not helping anything and
they're in the way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 7d734740be Addressing: avoid a Bfe
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 06299cca4b OpcodeDispatcher: optimize out Bfi for storereg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig af5aaab38b OpcodeDispatcher: optimize a few ALU-with-constant ops
instead of relying on ConstProp for this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 77bb01d384 OpcodeDispatcher: don't generate pointless Xor for AF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig a9eb1bd4e6 OpcodeDispatcher: don't generate pointless Bfe for moves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 83ef2da95f OpcodeDispatcher: avoid zero shift in SHLD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ba13ccadb8 OpcodeDispatcher: use ArithRef for 8/16-bit imul
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 4a2dee873d OpcodeDispatcher: use ArithRef for small rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 589906e6e6 OpcodeDispatcher: generalize AF on constants optimization
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig f0fcf6d9e6 OpcodeDispatcher: optimize SVE vmovmskpd
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 1591ced5a7 OpcodeDispatcher: use ArithRef for ADC/SBB flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 0d235d63d0 OpcodeDispatcher: use ArithRef for bit tests
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 820d5c1447 OpcodeDispatcher: use ArithRef for rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 8093e8f5f1 OpcodeDispatcher: optimize LoadDir
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 802857e411 OpcodeDispatcher: add ArithRef helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig bf3a1839a2 OpcodeDispatcher: transform DF ourselves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ad7844d7da OpcodeDispatcher: do not generate useless VMov
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 3ff9128a8a unittests: add asm test for mov ah, 0
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Ryan Houdek c865eb98ea Merge pull request #4598 from alyssarosenzweig/idc
Stop using hashmap in DCE
2025-06-02 10:26:08 -07:00
Ryan Houdek 217a9228f8 unittests/futimesat: Fixes flake
Due to interactions between file times and the lack of granularity in
futimesat, if we don't remove the nanoseconds then this test can flake.

Easy enough.
2025-06-02 10:21:42 -07:00
Alyssa Rosenzweig 105d0a36ad RedundantFlagCalculationElimination: dont use hashmap
combined results on node from this and the previous commit:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44433582    0.46988457    0.45344824    0.45320041  0.0043288623
+  50     0.4385365    0.46615359    0.45026179    0.44996216  0.0045742233
Difference at 95.0% confidence
	-0.00323825 +/- 0.00176704
	-0.71453% +/- 0.389903%
	(Student's t, pooled s = 0.00445323)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Alyssa Rosenzweig 263279d5dd IR: index blocks
this will let us avoid a costly hashmap in DCE.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Ryan Houdek 861ecbec0f Merge pull request #4600 from neobrain/fix_libfwd_cross_target
LibraryForwarding: Change target env to gnu
2025-06-02 09:21:46 -07:00
Ryan Houdek 5ff9bb3669 Merge pull request #4596 from pmatos/JSONOpts
Error when JSON config includes unknown options
2025-06-02 08:50:47 -07:00
Ryan Houdek e794584bb5 Merge pull request #4479 from neobrain/feature_codebuffer_sharing
Reduce JIT time by 25% by sharing code buffers between threads
2025-06-02 08:50:17 -07:00
Tony Wasserka a792dd0703 CPUBackend: Clean up CodeBuffer size limits 2025-06-01 22:45:50 +02:00
Tony Wasserka a9a6a645bf Arm64Emitter: Disable PC-relative constant encoding
This no longer works since the JIT output is now relocated before execution.
2025-06-01 22:45:50 +02:00
Tony Wasserka 7c93becd5f JIT: Increase estimate for CodeBuffer space use
The previous bound was exceeded during Steam startup before.
2025-06-01 22:45:50 +02:00
Tony Wasserka 4bbaef58e9 JIT: Re-enable parallel compilation by compiling to a temporary buffer 2025-06-01 22:45:50 +02:00
Tony Wasserka 95791a985a Core: Minimize the time CodeBufferWriteMutex is held 2025-06-01 22:45:50 +02:00
Tony Wasserka 0dfefe9730 Core: Re-check LookupCache before running compiler backend
This further reduces lock contention by skipping the backend phase in case
another thread raced the active one for the same block.
2025-06-01 22:45:49 +02:00
Tony Wasserka 8481c797df FEXCore/JIT: Extend LookupCache lock to all of ExitFunctionLink 2025-06-01 22:44:49 +02:00
Tony Wasserka a503b5e20b Rename CodeBufferManager reference 2025-06-01 22:44:49 +02:00
Tony Wasserka 4078840ef1 Core: Reduce JIT time by sharing CodeBuffers between threads
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
  be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
  in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
  the modifying thread. Old versions of the CodeBuffer persist as read-only
  data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
  decrease the reference count and eventually trigger deallocation of the
  old version
2025-06-01 22:44:49 +02:00
Tony Wasserka ab51958b26 Move AllocateNewCodeBuffer from CPUBackend to a new CodeBufferManager interface 2025-06-01 22:42:55 +02:00
Tony Wasserka 3376587b6a CPUBackend: Manage CodeBuffers using shared_ptr
This is required for sharing CodeBuffers between threads anyway, but it also
allows use of the constructor/destructor to manage memory automatically.
2025-06-01 22:42:55 +02:00
Tony Wasserka 6681d7dcf9 LookupCache: Split L3 cache into a dedicated interface
This data isn't really a cache, since the JIT is directly responsible of
writing its contents. Instead it be considered the source to populate the
L1/L2 caches from.

Furthermore, splitting off this data allows it to be shared across threads
in the future without affecting L1/L2 caches.
2025-06-01 22:42:55 +02:00
Tony Wasserka 34224481c6 LookupCache: Prefer empty() over a size check 2025-06-01 22:42:55 +02:00
Tony Wasserka 1837aaabe4 fextl: Add shared_ptr and make_shared 2025-06-01 22:42:55 +02:00
Tony Wasserka 5811914a78 LibraryForwarding: Change target env to gnu
This is required for clang to find architecture-specific libstdc++ headers as
distributed by NixOS.
2025-06-01 21:40:33 +02:00
Paulo Matos a109a4efa1 Error when JSON config includes unknown options 2025-06-01 11:59:51 +02:00
Alyssa Rosenzweig a08a6ce5de Merge pull request #4576 from Sonicadvance1/fix_vma_race
Linux/SMCTracking: Fixes nasty race condition causing invalid memory tracking
2025-05-30 08:14:38 -04:00
Ryan Houdek 3dc8a3ddc1 Merge pull request #4595 from neobrain/refactor_config_templates
Config: Clean up use of templates
2025-05-29 12:26:34 -07:00
Ryan Houdek 5e103365f7 Merge pull request #4594 from neobrain/fix_self_move
Async: Don't destruct on self-moves
2025-05-29 12:26:24 -07:00
Ryan Houdek cdea8d7f74 Merge pull request #4593 from alyssarosenzweig/ici/zeroing-sub-regs
InstructionCountCI: add more cases for mov 0/~0
2025-05-29 12:26:15 -07:00
Ryan Houdek ef6dc3d802 Merge pull request #4592 from alyssarosenzweig/opt/x87-tag
OpcodeDispatcher: optimize X87FTWTag
2025-05-29 12:26:04 -07:00
Ryan Houdek e6edb349ba Linux: Fixes some vestigial mmap handling
Missed this in previous commits, had to double check that I got them
all.
2025-05-29 12:15:26 -07:00
Ryan Houdek ad132267ec Linux/SMCTracking: Fixes nasty race condition causing invalid memory tracking
This has been a bug that we have technically lived with ever since SMC
tracking was introduced. The problem boils down to the fact that memory
management syscalls from multiple threads can race our SMC tracking.

This was only uncovered due to recent changes in the Steam client where
downloading games has more aggressively started reallocating memory.
This causes Steam to oversubscribe the CPU by a small margin, causing
threads to context switch more heavily during memory management.

The strace that finally managed to capture this:
```
41574 munmap(0xba84e000, 724992 <unfinished ...>
<...>
41227 mmap(NULL, 540672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -3, 0 <unfinished ...>
<...>
41574 <... munmap resumed>)             = 0
<...>
41227 <... mmap resumed>)               = 0xba87b000
```

While FEX's tracking linearly was:
```
mmap, 0xba87b000, 0x84000, 0x3, 0x22, 0xfffffffd, 0x0
munmap, 0xba84e000, 0xb1000
```

The way munmap and mmap perfectly interleave while getting context switched meant that the kernel's view of munmap then mmap didn't match our view of mmap completing first then munmap happening afterwards.
The kernel/strace is obviously the correct view in this instance.

This all comes down to how these threads are racing the VMA tracking
mutex after the syscall happens and not guaranteeing sequential
consistency that matches the kernel's view.

The only way to correct this sanely is to extend the locking period to
also encompass the syscalls getting executed. This is a bit tricky since
the VMA tracking needs to ensure that the lock is no longer held once
ThreadManager invalidation occurs so a callback to do the syscall
operation is about the only sane approach here. Luckily we now have
fextl::move_only_function.

Fixes consistent crashes with Steam game downloads (and maybe some
chromium crashes?)
2025-05-29 12:15:26 -07:00
Ryan Houdek 6f837281ef Merge pull request #4591 from Sonicadvance1/fix_warning2
FEXCore/CPUID: Remove warning
2025-05-29 11:21:49 -07:00
Tony Wasserka 2a1d29d2df Config: Clean up use of templates 2025-05-29 18:38:35 +02:00
Tony Wasserka bdceb4ca89 Async: Don't destruct on self-moves
FEXServer's logger performs such self-moves for the log pipe posix_descriptor
of short-lived clients. This resulted in a double-close previously, which
could interfere with file operations on other threads (typically crashing
FEXServer in effect).
2025-05-29 18:28:56 +02:00
Alyssa Rosenzweig 656bb928cf InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:53:38 -04:00
Alyssa Rosenzweig 0fbe69ebcf OpcodeDispatcher: optimize xor-with-self flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:53:38 -04:00
Alyssa Rosenzweig ece817c691 OpcodeDispatcher: optimize logical flags
seems to be strictly better.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:34:50 -04:00
Alyssa Rosenzweig 3fbc8204b7 OpcodeDispatcher: clean up logical flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:28:31 -04:00
Alyssa Rosenzweig 44a5481254 InstructionCountCI: add more cases for mov 0/~0
some obvious opportunities to improve here! although mostly I want this to
regression test my constprop rework.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:26:12 -04:00
Alyssa Rosenzweig cf2ff90f87 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:14:06 -04:00
Alyssa Rosenzweig b8dd5d95b0 OpcodeDispatcher: optimize X87FTWTag
using bit twiddling tricks :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:13:16 -04:00
Ryan Houdek 7cd52febc2 FEXCore/CPUID: Remove warning
This is currently only used on win32 builds because Linux doesn't
understand TPIDRRO.
2025-05-28 09:28:41 -07:00
Ryan Houdek 8e079c1965 Merge pull request #4590 from Sonicadvance1/remove_stack_set_arm
TestHarnessRunner: Stop setting the guest RSP for ARM test runner.
2025-05-27 09:21:14 -07:00
Ryan Houdek 4928af5a64 Merge pull request #4585 from alyssarosenzweig/opt/pair-push-pop
Pair push/pop
2025-05-27 09:17:47 -07:00
Alyssa Rosenzweig 289df740cd InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:33:13 -04:00
Alyssa Rosenzweig 6ad7392cd7 RegisterAllocationPass: pair push/pop
as a simple post-RA peephole. much much easier to do post-RA than pre-RA.

This isn't a post-RA /pass/ in the traditional sense... it's done while
assigning registers to coalesce the passes over the IR, since we pay per-pass
and we can merge the walks over the IR.

Closes: #4480
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:33:12 -04:00
Alyssa Rosenzweig b15d5f299c RegisterAllocationPass: drop dead IP increment
written but never read.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:32:33 -04:00
Ryan Houdek 5c7c959dd9 TestHarnessRunner: Stop setting the guest RSP for ARM test runner.
This matches behaviour with the x86 host runner that RSP isn't
guaranteed to be set to a valid memory location. Make sure when running
tests on ARM that it gets the same behaviour.
2025-05-26 13:02:07 -07:00
Alyssa Rosenzweig 6927c7577a RegisterAllocationPass: delete trivial instructions
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 88682a457a IR: add RemovePostRA helper
Remove blows up because of use tracking, but we can do a much simpler version
for post-RA and elide lots of checks from trying to make Remove more general.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 61ae53cc03 IR: add paired PushTwo/PopTwo helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 0e67f30103 OpcodeDispatcher: make Push do the right thing and use it
this both optimizes and bug-fixes pusha while deleting a snotton of code.

Closes: https://github.com/FEX-Emu/FEX/issues/4589
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig babd6e9a7b RegisterAllocationPass: allow Copy on the input IR
useful for pusha.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 76caa2c6e3 unittests: add pop-to-same-reg test
an earlier version of this PR failed this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Ryan Houdek dc9f8aa855 Merge pull request #4580 from alyssarosenzweig/ir/inline-ra
IR: Inline registers into the IR
2025-05-26 09:44:51 -07:00
Ryan Houdek ded8b3284a Merge pull request #4584 from neobrain/feature_assert_source_location
LogManager: Print source location when failing assertions
2025-05-24 01:15:11 -07:00
Tony Wasserka 2e24ee7a5f LogManager: Print source location when failing assertions 2025-05-24 09:35:11 +02:00
Tony Wasserka 048ae597d2 CodeEmitter: Fix incorrect macro parameter passing
Macros aren't aware of C++ templates, so the comma is considered a macro
argument separator unless the argument is wrapped in parentheses.
2025-05-24 09:35:11 +02:00
Ryan Houdek 0147f7aa19 Merge pull request #4587 from stanfordzhang/main
Fix callee saved floating-point arguments order issue
2025-05-23 20:13:58 -07:00
StanfordZhang 23cda2c961 Update Arm64Emitter.cpp
fix callee saved floating-point arguments issue
2025-05-23 22:14:37 +08:00
Alyssa Rosenzweig feb67658e1 RegisterAllocationPass: exploit new IR
Now that we can just set registers directly, we can simplify RA a lot. All the
Map/Unmap nonsense - it all goes away. We just assign registers as we go and
everything clicks into place naturally.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 65ee1fafa8 IR: drop RegisterAllocationData
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 49f8332c5b JIT: use registers directly from the IR
This is the flag day change from the series, using all the new shiny
infrastructre we added.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig afce108ed7 JIT: make almost all the DEF_OPs common
this deduplicates a bunch of #defines, letting us change the signature easier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 30f3b545af RegisterAllocationPass: use Header spill slots instead
Removes even more RAData dependence.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig bf1597920c IR: use post-RA flag
rather than implicitly depending on the RA data.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig f66bf3811e RegisterAllocationPass: set PostRA flag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig a01e29ac99 IR: extend the IR header with RA info
Beyond the actual registers allocated, there are two pieces of sideband data we
store in the RAData object:

* # of spill slots (explicitly)
* whether RA has run (implicitly by the existence of RAData)

We want to get rid of RAData, so we'll move these to the header.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 3136a5e2f8 IR: add helpers to extract physical registers from IR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig c4f00a05df IR: add space for registers right in OrderedNode *
This again follows the same idea of eliminating the RAData sideband.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig e9a8f8a9ff IR: generate builders that take OrderedNodeWrapper
when we need to materialize instructions inside RA without having a
corresponding OrderedNode* source, we want to just pass a OrderedNodeWrapper
with an encoded register. generate appropriate builders so this is possible.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 261ae7a195 IR: add immediates into OrderedNodeWrapper
this will let us encode registers directly inside OrderedNodeWrapper, rather
than pointers to OrderedNode *. that will let us speed up RA & onwards.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 7eaf5ae9e0 IR: drop FillRegister original source
this is now unused, and it's problematic with future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig 10a02449b1 IR: drop RA validation
There's no reasonable way to keep this around without adding significant
complexity to RA. This series prefers to drop complexity from RA, lessening the
need for validation in the first place.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig debc57e8c7 Merge pull request #4575 from alyssarosenzweig/cleanup/no-implicit-size
IR: drop DestSize inference
2025-05-16 15:17:53 -04:00
Tony Wasserka 7a4fff8e5b Merge pull request #4578 from cjacek/libgcc
Use -print-libgcc-file-name to get the compiler-rt file name
2025-05-16 13:40:25 +02:00
Jacek Caban ec0a2a8671 Use -print-libgcc-file-name to get the compiler-rt file name
Avoid hardcoding the file name. Upstream Clang currently expects the aarch64 variant of
compiler-rt. Changing this to use arm64ec in the file name could be problematic in the future,
if the MinGW toolchain gains support for ARM64X. Querying Clang for the correct file name is
the most forward-compatible approach.
2025-05-16 11:24:57 +02:00
LC b8b516f7b6 Merge pull request #4563 from Sonicadvance1/minor_vma_cleanup2
SyscallsVMATracking: Minor cleanup
2025-05-15 19:15:20 -04:00
Alyssa Rosenzweig 3b1e91b1fc IR: drop DestSize inference
no more users, and it's problematic for upcoming work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 15:00:03 -04:00
Alyssa Rosenzweig 8532593d91 IR: specify DestSize for RMWHandle
seems to just have been an oversight.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:58:05 -04:00
Alyssa Rosenzweig 55bd16e2c8 IR: make FillRegister sizes explicit
instead of hacking around it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:57:51 -04:00
Alyssa Rosenzweig c7797d56c9 IR: don't use GetOpSize in ExitFunction
nothing else does this, and it complicates upcoming refactor to move away from
IR builder helpers doing IR dereferencing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:51:11 -04:00
Ryan Houdek 9b573effd1 Merge pull request #4573 from alyssarosenzweig/cleanup/jit-id-2
JIT: use .ID() even less
2025-05-14 13:36:33 -07:00
Alyssa Rosenzweig dfd1aedae5 JIT: use .ID() even less
oops, missed a "/g" with th sed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 16:23:22 -04:00
Ryan Houdek 221ae2d7b4 Merge pull request #4571 from alyssarosenzweig/opt/constprop-xor-1
ConstProp: optimize XOR with all-1
2025-05-14 12:59:12 -07:00
Ryan Houdek 9127d206b5 Merge pull request #4572 from alyssarosenzweig/cleanup/jit-id
JIT: stop using .ID() pattern
2025-05-14 12:59:01 -07:00
Alyssa Rosenzweig ecc6fea54e JIT: stop using .ID() pattern
sed -ie 's/.ID()//' *

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:46:32 -04:00
Alyssa Rosenzweig 99446da7c1 JIT: add Reg helpers taking OrderedNodeWrappers
more ergonomic and will give us freedom to migrate things easier soon.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:43:20 -04:00
Alyssa Rosenzweig ad75563a26 IR: drop irrelevant reference to x86
we don't run the JIT on x86 anymore.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:32:07 -04:00
Alyssa Rosenzweig 197af972d9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 14:48:29 -04:00
Alyssa Rosenzweig fbd706c191 ConstProp: optimize XOR with all-1
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 14:48:29 -04:00
Ryan Houdek 6b85fa5611 SignalScopeGuards: Review 2025-05-14 10:56:44 -07:00
Ryan Houdek fae66a921f SyscallsVMATracking: Moves MappedResources implementation details to private 2025-05-14 10:56:44 -07:00
Ryan Houdek e6fc462e9d SyscallsVMATracking: Rename VMA tracking functions
Makes it more easily reflect what this functions are actually doing and
add a couple lines of documentation to help with the inherent opaqueness
of these.

NFC, just helps my brain wrap this more easily.
2025-05-14 10:56:44 -07:00
Ryan Houdek d11a265a53 SyscallsVMATracking: Adds validation that thread has ownership of VMA mutex
This way we can capture any programming bugs. Can only check the
functions that actually require unique locks rather than shared locks.
2025-05-14 10:56:44 -07:00
Ryan Houdek 395d870814 SignalScopeGuards: Add checks for locks being held by the calling thread
pthreads allows us to check if mutex/rwlock is currently locked by the
calling thread. This can give us some safety in code expecting locks to
be in place, allowing us to find programming bugs.
2025-05-14 10:56:44 -07:00
Ryan Houdek fec1ffaa6b Merge pull request #4566 from neobrain/feature_pool_alloc_size
ThreadPoolAllocator: Add support for updating the size of managed data
2025-05-14 10:53:34 -07:00
Tony Wasserka d4fdb28e72 PoolBufferWithTimedRetirement: Add support for updating the buffer size
This should be done at low frequency since it may unclaim the buffer.
2025-05-14 13:37:46 +02:00
Tony Wasserka 39a5c2021e ThreadPoolAllocator: Rename FixedSizePoolAllocation to PoolBufferWithTimedRetirement
This more accurately reflects that the core feature of the helper is the
timer-based unclaiming of buffers instead of the allocation size.
2025-05-14 13:29:44 +02:00
Tony Wasserka f41501444d ThreadPoolAllocator: Add a dedicated interface to try reowning a buffer without fallback 2025-05-14 13:29:44 +02:00
Tony Wasserka 86b26b80ce FixedSizePooledAllocation: Clean up documentation 2025-05-14 13:29:44 +02:00
Ryan Houdek ae9a5b1125 Merge pull request #4487 from pmatos/GCCTargetTestsN1
Run gcc target tests with block size 1
2025-05-13 04:45:52 -07:00
LC 38579807c2 Merge pull request #4567 from neobrain/fix_glibcxx_debug
ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
2025-05-09 15:17:48 -04:00
Tony Wasserka c19119bcd6 ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
Default-constructed iterators can't be copied.
2025-05-09 15:02:49 +02:00
LC 47f1ad693d Merge pull request #4561 from Sonicadvance1/minor_vma_cleanup
LinuxEmulation: Minor cleanup by separating VMA definitions
2025-05-08 15:21:44 -04:00
Ryan Houdek 2164d7bb96 Merge pull request #4565 from neobrain/fix_async_timeout
FEXServer: Don't time out while clients are still connected
2025-05-08 12:14:56 -07:00
Tony Wasserka c326e2d669 FEXServer: Don't time out while clients are still connected 2025-05-08 11:56:19 +02:00
Tony Wasserka 8eaf45414c Async: Add run_one interface to enable more fine-grained event loop control 2025-05-08 11:56:19 +02:00
Ryan Houdek b3297d106e Merge pull request #4564 from bylaws/arm64ec
CMake: Allow disabling explicit -mcpu usage
2025-05-07 15:30:04 -07:00
Ryan Houdek 2d0e19e6a7 Merge pull request #4562 from WhatAmISupposedToPutHere/main
Windows: Fix building with llvm-libcxx
2025-05-07 15:06:22 -07:00
Billy Laws fd2ee4dc46 CMake: Allow disabling explicit -mcpu usage 2025-05-07 22:48:05 +01:00
Sasha Finkelstein b1bbc37c59 Windows: Fix building with llvm-libcxx
Libcxx uses GetSystemTimePreciseAsFileTime if _WIN32_WINNT specifies
a version new enough to have it.
2025-05-07 23:47:24 +02:00
Ryan Houdek 56409d4f2b LinuxEmulation: Minor cleanup by separating VMA definitions
NFC, just moving this to its own header. It's already a huge PITA to
read. I just want to try and preserve some sanity while attempting to
fix #4557
2025-05-07 13:55:44 -07:00
Paulo Matos ec0683b729 Run gcc target tests with block size 1
info files for the test runner like Disabled_Tests, Known_Failures,
etc, receive not just the filename but the test name (which is the test name
and potentially its running config).

Increase the timeout of gcc tests to 30secs.
2025-05-07 16:14:53 +02:00
LC 89e5041e70 Merge pull request #4559 from Sonicadvance1/we_require_more_machicolations!
LinuxEmulation: Implement custom longjump that is fortification safe
2025-05-07 09:07:12 -04:00
LC 8b3e7312e0 Merge pull request #4560 from Sonicadvance1/fix_bad_define_check
LinuxEmulation: Fix bad compile time definition check
2025-05-06 22:28:41 -04:00
Ryan Houdek 3c76a9176d LinuxEmulation: Fix bad compile time definition check
We don't want this to be compiled out if the definition doesn't exist.
Actually define it in that case.
2025-05-06 18:39:58 -07:00
Ryan Houdek a37def2c22 LinuxEmulation: Implement custom longjump that is fortification safe
With fortifications enabled, glibc long jump has some additional checks
in place that break because we do a stack pivot. The only way around
this is to do our own long jumps. Luckily this is trivial.

Fixes #4558
2025-05-06 15:29:42 -07:00
Ryan Houdek ee4ef0fe87 Docs: Update for release FEX-2505 2025-05-05 11:37:17 -07:00
LC 1dff7073de Merge pull request #4554 from Sonicadvance1/remove_unused_argument
NFC: FEXCore: Removes unused argument on CreateThread
2025-05-05 14:30:24 -04:00
Ryan Houdek 92a82c3134 FEXCore: Removes unused argument on CreateThread
ParentTID is purely a Linux construct and has been moved entirely to the
frontend at this point. Remove this argument which is now unused.
2025-05-05 11:17:02 -07:00
Ryan Houdek a6a203c483 Merge pull request #4553 from neobrain/fix_align16b 2025-05-05 10:35:55 -07:00
Tony Wasserka cdaa65f6fd Arm64Emitter: Fix overalignment in Align16B
Previously, 16 additional bytes were emitted if the buffer was already
aligned.
2025-05-05 16:02:42 +02:00
Ryan Houdek 31e5f706b8 Merge pull request #4551 from OFFTKP/fadvise64
Fix 32-bit fadvise64
2025-05-02 18:37:55 -07:00
offtkp 00e05558ce Formatting 2025-05-03 04:28:23 +03:00
offtkp c02b88baff Fix 32-bit fadvise64 2025-05-03 03:56:16 +03:00
Ryan Houdek 794c80edad Merge pull request #4549 from OFFTKP/patch-2
Marshal freeram in sysinfo
2025-05-02 17:34:52 -07:00
Paris Oplopoios fb939600e2 Marshal freeram in sysinfo 2025-05-03 03:23:18 +03:00
Ryan Houdek 9dbbd44d09 Merge pull request #4544 from Sonicadvance1/fhu_ring_buffer
FHU: Add a non-block thread local ringbuffer
2025-05-02 11:22:39 -07:00
Ryan Houdek afcd93fe5f Merge pull request #4547 from Sonicadvance1/dont_go_chasing_sigsegv_waterfalls
Linux/SMCTracking: Stop calling mprotect on a memory region times the number of threads
2025-05-02 11:22:30 -07:00
Ryan Houdek 118faa5380 FHU: Add a non-block thread local ringbuffer
I keep rewriting this thing when I want to see some history on
something. Throw it in a FHU utility so I can stop wasting my time.
2025-05-02 01:19:00 -07:00
Ryan Houdek 7ed17f68c5 FEXCore: Remove unused InvalidateGuestCodeRange with callback 2025-05-02 01:15:50 -07:00
Ryan Houdek 96363143de Linux/SMCTracking: Stop calling mprotect on a memory region times the number of threads
I noticed this cascade of mprotects when poking at Crypt of the
Necrodancer, since it consistently is invalidating code. I saw us
calling mprotect on the same page 32 times in a tight loop and thought
surely this isn't FEX doing this.

Turns out we were calling the callback after invalidating each thread.
It should instead be done once at the end of invalidating the thread's
caches while still holding the locks.

Fixes this weird cascade of mprotects that equal the number of FEX
threads.
2025-05-01 21:14:05 -07:00
Ryan Houdek cf5fcb2371 Merge pull request #4546 from OFFTKP/shmdt
Make shmdt reset the unmapped pages in the 32-bit allocator
2025-05-01 20:17:50 -07:00
offtkp 7d3699050e Make shmdt reset the unmapped pages in the 32-bit allocator 2025-05-02 02:49:39 +03:00
LC 40e0a71f49 Merge pull request #4545 from Sonicadvance1/fix_fexserver_with_images
FEXServer: Fixes squashfs/erofs with newer fuse releases
2025-05-01 17:55:57 -04:00
Ryan Houdek cdbcdfe397 FEXServer: Fixes squashfs/erofs with newer fuse releases
Newer fuse releases changed how they are waiting on child processes to
exit. Setting the signal action to SIG_IGN would cause
erofsfuse/squashfuse to inherit the ignored action and cause their
internal `wait4` syscalls to fail with ECHLD.

Set the action back to default inside the FEXServer because our original
reasoning for setting the ignoring is no longer valid. FEXInterpreter
still ignores SIGCHLD while launching FEXServer.

Maybe fixes the muvm thing people have been complaining about.

Also fixes accidental comma delimiter usage.
2025-05-01 13:47:46 -07:00
LC 4f2d2e646e Merge pull request #4542 from Sonicadvance1/fexcore_reconstructions_getting_saved_today
FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
2025-05-01 13:28:16 -04:00
LC f97cd24647 Merge pull request #4543 from Sonicadvance1/fix_my_reducing_of_x87_today
FEXCore: Fixes x87 reduced precision
2025-05-01 13:25:30 -04:00
Ryan Houdek e5e75ad1ef InstcountCI: Update 2025-04-30 17:54:55 -07:00
Ryan Houdek fc052efb91 FEXCore: Fixes x87 reduced precision
With the change from #4538 I had accidentally broken x87 reduced
precision.

This is due to the fact that we accidentally lost ABI information about
interpreter fallbacks supporting `preserve_all` or not. So now instead
of having some ABI callbacks supporting it and some not, just force
usage of `preserve_all` if it is supported by the compiler entirely.

Fixes Steam when x87 reduced precision is enabled.
2025-04-30 17:37:10 -07:00
Ryan Houdek c5754145c5 FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
As said in the implementation of this struct commit message. This new
pair struct optimizes specific cases of small forward only increments
that can fit in to 8-bit space, and small forward or backward jump cases
that fit in to 16-bit space.

Some stats of this change:
- Steam: 5.88MB down to 4.34MB. 73.8% the space consumed
- Steamwebhelper: 15.8MB down to 13.24MB. 83.8% space consumed
- Sonic Mania: 3.6MB down to 2.58MB. 71.6% space consumed

As for absolute stats when compared to all code buffer size:
- Steam: 86MB of code buffer to 5.88MB -> 4.34MB of RIP reconstruction.
  - 6.8% -> 5% code buffer space used for RIP reconstruction
- Steamwebhelper: 285MB of code buffer to 17MB -> 14.26MB of RIP reconstruction.
  - 5.9% -> 4.9% code buffer space used for RIP reconstruction
- Sonic Mania: 48.53MB of code buffer to 3.55MB -> 2.53MB of RIP reconstruction.
  - 7.3% -> 5.2% code buffer space used for RIP reconstruction
2025-04-30 15:41:43 -07:00
Ryan Houdek 9986e69622 FEXCore/APITests: Extend tests for vl64pair 2025-04-30 15:41:20 -07:00
Ryan Houdek 00acf4d327 FEXCore/Utils: Implement a new VL64Pair struct type
This new variable length pair of integer implementation is taking direct
advantage of the most common aspects of FEX's JIT in that most x86
instructions are <= 8-bytes in length, and the ARM implementations of
those are /usually/ 16 instructions in length or less. Also only
unsigned offsets in this implementation since the common case is forward
incrementing.

This converts a majority of 16-bit vl encodings in to an 8-bit encoding
instead, shaving space off the RIP reconstruction data.

An additional optimization is for the 16-bit pair of integers, we
continue this optimization through but with more bits and changing over
to signed. This captures the second most common cases of /slightly/
larger increments and small loops.

Pairs of 32-bit and 64-bit integers are unoptimized since they are
uncommon.
2025-04-30 15:36:49 -07:00
Ryan Houdek 00ff549044 Merge pull request #4538 from Sonicadvance1/fex_interpreters_are_dancing
JIT: Move interpreter ABI handlers in to the dispatcher
2025-04-30 11:51:29 -07:00
Ryan Houdek 75a39c8939 InstcountCI: Update 2025-04-29 22:35:39 -07:00
Ryan Houdek 5b9354cc74 JIT: Move interpreter ABI handlers in to the dispatcher
Fixes #4535

A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher

This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.

This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
2025-04-29 22:35:39 -07:00
Ryan Houdek f7049a6478 Arm64Emitter: On spill return stack used and stop clobbering TMP4
TMP4 was used before we passed in a tmp register. Now use that temp
register.

Also return the amount of stack used on the push function. This will be
used in a bit.
2025-04-29 22:31:35 -07:00
Ryan Houdek e1d032b5a6 Merge pull request #4540 from pmatos/MProtectLastPage
mprotect last page of CodeBuffer
2025-04-27 11:07:14 -07:00
LC f99691b0eb Merge pull request #4541 from Sonicadvance1/f80_cephes_softfloat_prep
80-bit cephes prep work
2025-04-26 09:18:57 -04:00
Ryan Houdek 10d18f2f17 Softfloat-3e: Adds missing extF80_le file 2025-04-25 17:30:18 -07:00
Ryan Houdek a236700c55 cephes_128bit: Rename a bunch of variables
These are going to conflict with the 80-bit implementation otherwise.
2025-04-25 17:30:18 -07:00
Ryan Houdek 34d62fcea6 cephes: Split 128-bit to its own folder
We are soon going to learn how to operate in f80.
2025-04-25 17:30:18 -07:00
Paulo Matos 791502afef Protect last page of CodeBuffer
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.

In addition, there a couple of clang-tidy fixes which should be NFC.
2025-04-25 20:41:18 +02:00
Ryan Houdek 99114e1fc2 Merge pull request #4528 from neobrain/feature_unlink_threadsafe
JIT: Make code patching during (un-/)linking thread-safe
2025-04-25 08:22:38 -07:00
Ryan Houdek 3c8bb53d1a Merge pull request #4539 from cjacek/mincore
Link to mincore instead of kernelbase on ARM64EC
2025-04-25 08:22:19 -07:00
Jacek Caban bd0682b518 Link to mincore instead of kernelbase on ARM64EC
The kernelbase import library is not available in upstream mingw-w64. Instead, similar to MSVC,
we can use mincore, which allows importing the relevant functions via apisets.
2025-04-25 11:38:16 +02:00
Tony Wasserka e89a913b69 JIT: Make memory write visible to other threads reading the same location 2025-04-25 09:47:13 +02:00
Tony Wasserka 293d77d412 JIT: Make code patching during (un-/)linking thread-safe 2025-04-25 09:21:09 +02:00
Ryan Houdek 6beb4b0f8b Merge pull request #4533 from Sonicadvance1/align_tail_in_the_pale_moonlight
JIT: Align JITCodeTail to native alignment
2025-04-24 16:18:45 -07:00
LC 23a8462688 Merge pull request #4536 from Sonicadvance1/eternal_sunshine_of_a_vectorless_mind
JIT: Moves VPCMPESTRX handler to use vectors
2025-04-24 13:35:33 -04:00
LC 02f90fbc66 Merge pull request #4534 from Sonicadvance1/spaaaaaaace
CodeEmitter: Fix clang-format
2025-04-24 13:34:19 -04:00
Ryan Houdek 8cc23d0cf8 Merge pull request #4537 from sdpoueme/main
updated Dockerfile to reflect latest FEX-Emu releases
2025-04-24 10:28:24 -07:00
Serge Poueme 60b5b1414d updated Dockerfile to reflect latest FEX-Emu releases
Add multi-stage Dockerfile for FEX emulator build

- Stage 1 (Builder): Sets up build environment with Ubuntu 22.04
  - Installs development tools and dependencies
  - Builds FEX using clang-13 with optimized settings
  - Configures cmake with LTO enabled and tests disabled

- Stage 2 (Runner): Creates minimal runtime image
  - Includes only necessary runtime libraries
  - Copies built binaries from builder stage
2025-04-24 10:13:04 -07:00
Ryan Houdek cbab344822 InstcountCI: Update 2025-04-24 10:11:25 -07:00
Ryan Houdek 32764ddf81 JIT: Moves VPCMPESTRX handler to use vectors
I pushed this off from the previous changes that were converting things
to vector as less important. It has now become more important to keep
these in vector registers until beyond the ABI boundary.

This will reduce burden on our JIT backend and just changes where the
movement in to GPRs occurs. Necessary for #4535
2025-04-24 10:09:26 -07:00
Ryan Houdek e31127e39e CodeEmitter: Fix clang-format 2025-04-24 08:41:23 -07:00
Ryan Houdek 735f537846 JIT: Align JITCodeTail to native alignment
Removes UB
2025-04-24 08:40:43 -07:00
Ryan Houdek b7790e10e9 Merge pull request #4532 from pmatos/X87StateBlockReset
X87 state block reset
2025-04-24 08:37:32 -07:00
Ryan Houdek 3f7ad04054 Merge pull request #4531 from Sonicadvance1/instcountci_change_size
InstcountCI: Changes how code size is calculated
2025-04-24 08:35:33 -07:00
Paulo Matos 670ddf9cd8 instcountci: Reset MMXState to X87 at the start of each block 2025-04-24 15:02:46 +02:00
Paulo Matos d377e26106 Reset MMXState to X87 at the start of each block
Ensures that blocks always start with the same state independently of predecessors
which allows independent compilation of blocks.
Starting in the X87 state is better than starting in MMX state because
MMX state is more work to initialize.
2025-04-24 15:02:41 +02:00
Ryan Houdek 674f939690 Merge pull request #4523 from bylaws/maptrack
Improve tracking of executable mappings
2025-04-23 13:08:12 -07:00
Ryan Houdek 64abfb1afe InstcountCI: Update 2025-04-23 12:56:46 -07:00
Ryan Houdek adcefe62cf InstcountCI: Changes how code size is calculated
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.

Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.

So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.

This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.

This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.

This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
2025-04-23 12:56:46 -07:00
Ryan Houdek bf77fa4275 unittests/Emitter: Adds udf test 2025-04-23 12:53:36 -07:00
Ryan Houdek e7fce2de40 CodeEmitter: Adds udf support 2025-04-23 12:14:44 -07:00
Ryan Houdek 8864637602 Merge pull request #4465 from Sonicadvance1/ubsan_fixes
Various: UBSAN fixes around unaligned accesses
2025-04-23 00:54:49 -07:00
Ryan Houdek 1a0d97d5f7 Various: UBSAN fixes around unaligned accesses 2025-04-23 00:31:09 -07:00
Ryan Houdek 9ba46e395d Merge pull request #4530 from neobrain/fix_instcountci_script
Scripts/InstCountCI: Restrict git-add to relevant JSON files only
2025-04-22 06:10:43 -07:00
Tony Wasserka be99fa0b05 Scripts/InstCountCI: Restrict git-add to relevant JSON files only 2025-04-22 13:34:28 +02:00
Ryan Houdek c42b1858c5 Merge pull request #4529 from lioncash/tests
ASIMD_Tests: Enable PMULL/PMULL2 tests
2025-04-21 19:03:01 -07:00
Lioncache cc11de16bf ASIMD_Tests: Enable PMULL/PMULL2 tests
vixl now supports these. We can also get rid of the invalid data sizes,
since only halfword and 128-bit variants are defined.
2025-04-21 20:03:37 -04:00
Ryan Houdek 3ea7c2b014 Merge pull request #4522 from bylaws/persona
LinuxEmulation: Copy the host persona on thread startup
2025-04-21 00:59:01 -07:00
Ryan Houdek e0e9f5aa16 Merge pull request #4527 from lioncash/concept
CodeEmitter/ASIMDOps: Constrain Q and D register requirements with concept
2025-04-20 18:14:15 -07:00
Lioncache 63a2d1e1db CodeEmitter/ASIMDOps: Move three emitter helpers to private section
These don't need to be public.
2025-04-20 15:55:53 -04:00
Lioncache 644263a764 CodeEmitter/ASIMDOps: Constrain Q and D register requirements with concept
Pulls these out into a single concept instead of having the same lengthy
requirements clause.
2025-04-20 15:55:49 -04:00
Ryan Houdek f912295690 Merge pull request #4524 from bylaws/faulto
X86Tables: Set FLAGS_BLOCK_END for more faulting ops
2025-04-19 19:02:09 -07:00
Ryan Houdek 79e57c320c Merge pull request #4525 from lioncash/pac
CodeEmitter/LoadstoreOps: Add Load/store register (PAC) group
2025-04-19 18:29:18 -07:00
Lioncache 4e6f183f77 CodeEmitter/LoadstoreOps: Add Load/store register (PAC) group
Eh, what the heck. Gets rid of the last TODO marker in the base load-stores.
2025-04-19 13:00:00 -04:00
Billy Laws 23017af874 TestHarnessRunner: Call the frontend mapping callback for created mappings 2025-04-19 15:29:23 +01:00
LC 4d02126b77 Merge pull request #4517 from Sonicadvance1/ruining_another_set_of_memory_leaks
LinuxEmulation: Fixes remaining memory leak on pthread teardown
2025-04-19 09:44:26 -04:00
Billy Laws 30fefbd57d X86Tables: Set FLAGS_BLOCK_END for more faulting ops 2025-04-19 14:19:19 +01:00
Billy Laws 38304fa63a LinuxEmulation: Copy the host persona on thread startup 2025-04-19 14:10:55 +01:00
Billy Laws 6525b8dd14 VDSO_Emulation: Map the VDSO thunk as executable 2025-04-19 14:09:48 +01:00
Billy Laws 0155719b1a VDSO_Emulation: Ensure the 32 bit sigreturn mapping is tracked as executable 2025-04-19 14:09:48 +01:00
Billy Laws 966dce6e95 SyscallsSMCTracking: Handle shmat SHM_EXEC flag 2025-04-19 14:09:48 +01:00
Billy Laws e3088686db InvalidationTracker: Track executable mappings 2025-04-19 14:09:48 +01:00
Billy Laws 771162cf14 InvalidationTracker: Track the protection of regions mapped at startup
Code can be injected into a process by e.g. the chromium sandbox before
FEX is loaded.
2025-04-19 14:09:48 +01:00
Billy Laws 1a2398d2ce Windows: Call the protection callback for executable FEX mappings 2025-04-19 14:09:48 +01:00
Billy Laws 0f0e969d8f Windows: Call the image map callback for ntdll 2025-04-19 14:09:48 +01:00
Ryan Houdek c8371087e6 Merge pull request #4519 from neobrain/fix_fexconfig_layout
FEXConfig: Fix layout issues on Qt 6.9
2025-04-18 09:25:08 -07:00
Ryan Houdek 3da6fc3972 Merge pull request #4518 from neobrain/refactor_warn_fixes
JIT: Fix warning about unused variable
2025-04-18 09:24:49 -07:00
Ryan Houdek 0cbbd91e72 Merge pull request #4521 from lioncash/mem
CodeEmitter/LoadStoreOps: Add Memory Copy and Memory Set category
2025-04-18 09:24:24 -07:00
Lioncache d71aca9e9b CodeEmitter/LoadStoreOps: Add Memory Copy and Memory Set category
Adds all of the memory facilities in FEAT_MOPS to the emitter.
2025-04-18 10:53:38 -04:00
Tony Wasserka 63035fd5f3 FEXConfig: Fix layout issues on Qt 6.9 2025-04-18 11:17:51 +02:00
Tony Wasserka 0fe28129a0 JIT: Fix warning about unused variable 2025-04-18 11:16:46 +02:00
Ryan Houdek c44d8eed35 LinuxEmulation: Fixes remaining memory leak on pthread teardown
With the previous stack leak fix, RUINER reduced its memory leaking down
to around 50MB/s. The remaining memory leaks come from the pthread stack
that we are required to allocate (128KB per thread) and some internal
DTV tracking structures.

The problems come in the fact that glibc/pthread only tears down its
internal state for these if the pthread function actually returns! We
can **technically** switch the initial stack over to a "user" stack but
that introduces more problems around internal dtv tracking that we
already fixed months ago, so we can't actually do that in practice.

This leaves us no choice, we effectively are mandated to return from the
pthread function in order to free the memory from glibc. The only way we
can safely do this is with a long jump and deferring some data structure
management until that case.

This is all incredibly sucky but it's necessary to work. With these
changes, RUINER is no longer leaking memory (Hovering at around 3GB used
while in-game) and even Steam is consuming less memory.

It doesn't solve the problem that thread creation and teardown could
likely be faster, but not many applications are creating 720
threads/second.
2025-04-17 18:16:32 -07:00
Ryan Houdek 6f5588f71e Merge pull request #4504 from pmatos/NoStrictAliasing
Enable -fno-strict-aliasing
2025-04-17 17:27:33 -07:00
Ryan Houdek 68f8c244b4 Merge pull request #4516 from lioncash/radd
CodeEmitter/ASIMDOps: Support RADDHN{2}/RSUBHN{2}
2025-04-17 17:27:22 -07:00
Lioncache d7a08fc7d4 CodeEmitter/ASIMDOps: Support RADDHN{2}/RSUBHN{2}
These are trivial enough to just drop right in.
2025-04-17 20:13:30 -04:00
LC 572d6e0395 Merge pull request #4509 from Sonicadvance1/in_the_twilight_of_the_pale_blue_moon
A couple of barrier and timing fixes.
2025-04-17 17:30:03 -04:00
LC 156b6745a2 Merge pull request #4513 from Sonicadvance1/fix_stack_leak
LinuxSyscalls: Fixes a major stack memory leak
2025-04-17 17:27:15 -04:00
Ryan Houdek b4eb38e2b7 Merge pull request #4515 from lioncash/scalar
CodeEmitter/ScalarOps: Add two more instruction categories
2025-04-17 14:10:04 -07:00
Lioncache bafcc8762b CodeEmitter/ScalarOps: Handle ASIMD scalar x indexed element group 2025-04-17 14:47:38 -04:00
Lioncache ec6ff7899d CodeEmitter/ScalarOps: Remove implemented TODO
These scalar ops are already implemented and this was accidentally left in.
2025-04-17 08:14:10 -04:00
Lioncache 6330b398a5 CodeEmitter/ScalarOps: Handle ASIMD scalar three same extra group
These are trivial enough to drop in.
2025-04-17 08:09:42 -04:00
Ryan Houdek 86c492da13 LinuxSyscalls: Fixes a major stack memory leak
Fixes an issue where a thread that exits with the `exit` syscall never
actually frees its pivot stack. This is common practice and it was
missed when I was fixing the previous stack pivot leak.

This was uncovered when looking at the game
[RUINER](https://store.steampowered.com/agecheck/app/464060/?curator_clanid=4777282)
for timing bugs. Turns out the Linux build of the game creates and
destroys a VLC object every tick of the engine. Creating this VLC
object creates six threads behind the scenes. At what I assume the
default tick of the engine is of 120Hz(?) this would mean it is
attempting to create and destroy 720 threads per second, quickly
leading to memory exhaustion under FEX.

While this hits the biggest memory leak we have around thread creation,
this game is still hitting thread creation so hard that I can see other
leaks that I need to track down still.
2025-04-16 19:47:56 -07:00
Ryan Houdek 8ff7497d59 Merge pull request #4501 from JunChi1022/mask_signal_at_defer
LinuxSyscalls: Update signal mask at deferring time
2025-04-16 10:56:40 -07:00
Ryan Houdek 73593ad4f1 Merge pull request #4499 from bylaws/x87fix
OpcodeDispatcher: Mark NZCV as dirty when clobbering in FCOMIF64
2025-04-16 10:32:13 -07:00
Ryan Houdek 572a58ed27 Merge pull request #4512 from lioncash/internal
Arm64: Mark several functions as internally linked
2025-04-16 10:31:06 -07:00
Ryan Houdek d7df4be379 Merge pull request #4511 from lioncash/test
ASIMD_Tests: Re-enable tests disabled due to dissassembly bugs
2025-04-16 10:30:25 -07:00
Ryan Houdek f227722898 Merge pull request #4503 from pmatos/UBSANflags
Add no-sanitize flags to UBSAN
2025-04-16 10:29:02 -07:00
JustinChi f371bf2bc4 LinuxSyscalls: Update signal mask at deferring time
In regular x86 programs, when a signal occurs, the signal will not be handled within the signal handler. However, under FEX's defer signal mechanism, the signal is not immediately masked when it is deferred. When returning to the location that receives the signal and continues processing, the signal might be received again, causing inconsistency between the emulation and the actual program.

Here is an unit test for this patch from ltp:
https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/syscalls/timer_settime/timer_settime03.c
2025-04-16 21:39:53 +08:00
Lioncache e6841ea46d Arm64_stubs: Remove non-public functions
These aren't exposed in the public interface anymore, so they can be removed.
2025-04-16 09:21:51 -04:00
Lioncache edbfe45a20 Arm64: Mark several functions as internally linked
These aren't directly used outside of the translation unit.
2025-04-16 09:16:44 -04:00
Lioncache 5f83e89be5 ASIMD_Tests: Re-enable tests disabled due to dissassembly bugs
These issues seem to be resolved now.
2025-04-16 08:47:18 -04:00
Billy Laws 5850b26de5 Update InstCountCI 2025-04-16 13:16:07 +01:00
Billy Laws 416267a238 OpcodeDispatcher: Safely clobber NZCV in FCOMIF64
Also fix a small typo that broke the !flagm2 path.
2025-04-16 13:06:37 +01:00
LC 94499ed8fc Merge pull request #4510 from Sonicadvance1/fix_space
CodeEmitter: Fixes misaligned function
2025-04-15 19:58:40 -04:00
Ryan Houdek da1868288d CodeEmitter: Fixes misaligned function 2025-04-15 15:47:31 -07:00
Ryan Houdek c8d6b39585 EmulatedFiles: Match bogomips calculation
Instead of hardcoding the bogomips calculation, more closely match what
the Linux kernel does for bogomips. There are some applications out
there that use bogomips for silly timing calculations, so this is a
better version.
2025-04-15 15:42:52 -07:00
Ryan Houdek 00f8181c3a OpcodeDispatcher: Fix CPUID being a instruction fence
It is common practice for games to use CPUID as an instruction barrier
for various reasons. Ensure that we respect this by adding support for
an instruction barrier.
2025-04-15 15:42:43 -07:00
Ryan Houdek 60565cc2ef OpcodeDispatcher: Implement support for self-synchronizing cycle counter
FEAT_ECV added a new synchronizing cycle counter instruction that
restrict speculation across the cycle counter access. Because it
restricts speculation, it effectively acts like an isb and load dsb.

Luckily for us, this actually matches behaviour for what rdtscp does, so
we can take advantage of it if the host supports FEAT_ECV.
2025-04-15 15:42:43 -07:00
Ryan Houdek bfee39ae70 HostFeatures: Passthrough if the host supports ECV 2025-04-15 15:38:00 -07:00
Ryan Houdek 753c5e72be CodeEmitter: Implement support for CNTVCTSS_EL0 2025-04-15 15:37:02 -07:00
Ryan Houdek 571a533543 Merge pull request #4508 from lioncash/crypto
Emitter: Add missing ASIMD crypto operations
2025-04-15 11:24:51 -07:00
Lioncache 222fcfd70e CodeEmitter/ASIMDOps: Add crypto two-register SHA512 category 2025-04-15 11:48:30 -04:00
Lioncache 95c8a70ce6 CodeEmitter/ASIMDOps: Add crypto four-reg category 2025-04-15 11:41:07 -04:00
Lioncache 0f6bec4d5e CodeEmitter/ASIMDOps: Add crypto three-reg SHA512 category 2025-04-15 11:31:00 -04:00
Lioncache 5fb0522c78 CodeEmitter/ASIMDOps: Add crypto three-reg imm2 category 2025-04-15 11:15:21 -04:00
Tony Wasserka 7c0bc2d972 Merge pull request #4506 from pmatos/CastFixValidateCode
Fix cast in ValidateCode impl
2025-04-15 14:07:51 +02:00
Paulo Matos cb972e165b Fix cast in ValidateCode impl 2025-04-15 09:35:46 +02:00
Paulo Matos fb319a352a Enable -fno-strict-aliasing 2025-04-14 09:29:15 +02:00
Paulo Matos e10646b436 Add no-sanitize flags to UBSAN
See discussion in https://github.com/FEX-Emu/FEX/pull/4494 for context.
2025-04-14 09:10:28 +02:00
Ryan Houdek 211bec65d2 Merge pull request #4492 from bylaws/badencodings
Frontend: Be more tolerant of bad instruction encodings
2025-04-13 19:09:07 -07:00
Ryan Houdek 66807539ce Merge pull request #4500 from bylaws/arm64ec-fixes-etc
Windows: Misc cleanups and fixes
2025-04-13 19:03:17 -07:00
Ryan Houdek 2919c32a8f Merge pull request #4497 from pmatos/CleanupGCCTests
Cleanup info test files for 32bits
2025-04-11 17:43:09 -07:00
Billy Laws 05be9e9c37 Windows: Process cross-process notifications before code compilation 2025-04-11 12:07:16 +01:00
Billy Laws 039aaae041 FEXCore: Add a pre-compilation frontend callback to SyscallHandler 2025-04-11 12:07:16 +01:00
Billy Laws 890653ae39 WOW64: Handle the image mapped callback 2025-04-11 12:07:16 +01:00
Billy Laws 8644fdd595 ARM64EC: Slight cleanup 2025-04-11 12:07:16 +01:00
Billy Laws 4fca8fb8e6 AllocatorHooks: Correctly restore the region protection in VirtualDontNeed 2025-04-11 12:07:16 +01:00
Billy Laws 9625201cbf Frontend: Remove asserts on invalid instruction encodings 2025-04-11 12:06:01 +01:00
Paulo Matos e7dea56adf NFC: Cleanup info test files for 32bits
No point in duplicating test info between Disabled and Known Failures.

Leaving tests known to fail in Known_Failures. Flakes / Unreliable tests go into Disabled_Tests.
2025-04-11 10:50:41 +02:00
LC 5b802d17d1 Merge pull request #4498 from Sonicadvance1/fix_llseek
LinuxSyscalls: Fixes 32-bit llseek result
2025-04-10 16:18:54 -04:00
Ryan Houdek a43ddba87c LinuxSyscalls: Fixes 32-bit llseek result
`llseek` returns only ever 0 or errno in the return register. This is in
contrast to `lseek` which returns the result (or errno) in the return
register.

We were accidentally returning the result on non-error conditions which
could freak out some software. Thanks to
[OFFTKP](https://github.com/OFFTKP) for pointing out this issue
2025-04-10 13:08:06 -07:00
Tony Wasserka f1d007ca4c Merge pull request #4495 from pmatos/CleanupCompWarn
Cleanup compile warnings
2025-04-10 16:40:46 +02:00
Paulo Matos c2111e7384 Cleanup compile warnings 2025-04-10 08:49:51 +02:00
Billy Laws a1c6378317 Frontend: Only accept POP opcodes with a 0 ModRM.reg field 2025-04-09 23:04:11 +01:00
Billy Laws f064013b9a Frontend: Enforce FLAGS_SF_MOD_MEM_ONLY/FLAGS_SF_MOD_REG_ONLY 2025-04-09 23:02:09 +01:00
Billy Laws d4480c3566 X86Tables: Mark VMOVNTDQ as FLAGS_SF_MOD_MEM_ONLY 2025-04-09 23:01:26 +01:00
Billy Laws b3a4de7aa8 X86Tables: Mark MASKMOVQ as FLAGS_SF_MOD_REG_ONLY 2025-04-09 23:01:26 +01:00
Billy Laws 60d8131a4d X86Tables: Drop FLAGS_SF_MOD_MEM_ONLY from (V)MOV(L/H)PS
These encodings are shared with MOVLHPS/MOVHLPS and FEX handles both variants.
2025-04-09 23:01:26 +01:00
Billy Laws 6d78aefa13 Frontend: Ignore REX register extension for MMX registers 2025-04-09 23:01:26 +01:00
Billy Laws badd953600 unittests: Test for REX.B being ignored with MMX registers 2025-04-09 23:01:26 +01:00
LC 5d385c540d Merge pull request #4489 from Sonicadvance1/reduce_stack_cpu_count
FEXCore: Reduce stack usage in CalculateNumberOfCPUs
2025-04-09 12:40:52 -04:00
Ryan Houdek 52351902ee Merge pull request #4488 from pmatos/EnableUBSAN
Add option to enable UBSAN
2025-04-09 00:13:29 -07:00
Paulo Matos 6e697e5ec1 Add option to enable UBSAN 2025-04-09 08:58:24 +02:00
Ryan Houdek c55c8979f1 AOTGen: Only calculate number of CPU cores once
Removes some per-iteration string processing and file IO.
2025-04-08 22:55:55 -07:00
Ryan Houdek 7ca757bb6d FEXCore: Move CPUInfo to FEX
This is only ever used in the frontend now.
2025-04-08 22:54:43 -07:00
Ryan Houdek c2d59b02bd FEXCore: Reduce stack usage in CalculateNumberOfCPUs
Nothing crazy, just recalculate the maximum string length rather than
use PATH_MAX.
2025-04-08 22:54:43 -07:00
LC 45b638fc45 Merge pull request #4491 from Sonicadvance1/push_pop_callee
FEXCore/Emitter: Stop creating a vector on the heap
2025-04-09 01:46:34 -04:00
Ryan Houdek 5767c61a91 FEXCore/Emitter: Stop creating a vector on the heap
In the Push/Pop CalleeSavedRegisters these vectors were getting created
on the heap, allocating memory and then just iterating them.

Just use a std::array which makes it stop allocating memory and saves
the number of instructions.
2025-04-08 18:36:47 -07:00
LC e3f8a817e0 Merge pull request #4490 from Sonicadvance1/remove_old_workaround
FEXCore/Allocator: Removes old workaround for kernel 4.17
2025-04-08 21:34:47 -04:00
Ryan Houdek 4377603d7e FEXCore/Allocator: Removes old workaround for kernel 4.17
This was only used for working around our old CI machines and now that
our minimum kernel requirement is 5.15 this isn't required anymore.
2025-04-08 18:22:39 -07:00
LC bc72239181 Merge pull request #4485 from Sonicadvance1/softfloat-3e_potara_cephes
cephes: Rewrite to use softfloat-3e 128-bit
2025-04-08 11:26:01 -04:00
Ryan Houdek 2d7f37386e Softfloat: Remove warnings
These precision warnings are no longer true!
2025-04-07 15:22:24 -07:00
Ryan Houdek dffba2d83c unittests/ASM: Enable x87 tests under simulator
Since the x86 simulator is now handling these instructions at 128-bit
precision, these now just work.
2025-04-07 15:13:04 -07:00
Ryan Houdek 0fbb6aa02f cephes: Rewrite to use softfloat-3e 128-bit
And also use it at the same time, since the function signatures changed.

Instead of relying on the host libc math libraries for `long double` ALU
operations, rewrite the entire thing to use softfloat-3e fixed width
float128_t types.

This is a very invasive change in cephes but is a necessary requirement
for getting the precision we require in environments that map `long
  double` to be the same as `double`, like Win32 and MacOS.

This fixes the precision issue in transcendental operations when running
under WINE.
2025-04-07 15:13:04 -07:00
Ryan Houdek 7364affd69 Softfloat-3e: Adds some missing files and passes state through 2025-04-07 14:57:48 -07:00
LC 09e622d5a0 Merge pull request #4482 from Sonicadvance1/softfloat_3e
Softfloat-3e: Add support for f128
2025-04-06 18:03:35 -04:00
Ryan Houdek 21ccaf56e6 Softfloat-3e: Add support for f128
This is going to be necessary soon.
2025-04-05 17:48:57 -07:00
Ryan Houdek d8cd807520 Softfloat-3e: Moves to Externals 2025-04-05 17:26:15 -07:00
LC 6069d05aa9 Merge pull request #4481 from Sonicadvance1/instcountci_ensure_consistent_rip
InstcountCI: Ensure RIP of blocks is consistent
2025-04-05 19:04:39 -04:00
Ryan Houdek 09a3d4851a InstcountCI: Update 2025-04-05 15:55:08 -07:00
Ryan Houdek d41374c55a InstcountCI: Ensure RIP of blocks is consistent
This changes the instcountCI code to consistently load test data in to
RIP 0x1'0000 so we don't have any spurious changes due to json changes.

This has been a minor annoyance where if a test was added, it had the
potential to shift the rest of the data in the tests. This now ensures
it is consistent.
2025-04-05 15:55:08 -07:00
LC dd8a1f00fc Merge pull request #4468 from Sonicadvance1/thats_the_softfloat_guarantee
FEXCore/Softfloat: Wire up cephes math library for transcendental operations
2025-04-04 19:11:55 -04:00
Ryan Houdek ffca27cbde FEXCore/Softfloat: Wire up cephes math library for transcendental operations
This is solving a different problem than what #4411 is specifically
trying to solve.

For our transcendental operations, we can't currently guarantee that
these functions will actually operate at the 128-bit softfloat
precision. While this is true with glibc, this is /not/ true for musl
and likely more libraries.

Instead of relying on our libc implementation to implement these,
instead include the cephes math library directly which is what most
people use for this. Including musl even, but not for all operations.

With this we are no longer beholden to the standard libraries for
providing a correct implementation.
2025-04-04 16:03:09 -07:00
Ryan Houdek 82edd901fe FEXCore/Common: Adds cephes math library
Only the few transcendental functions that FEX needs.
Disabled when building on x86-64.
2025-04-04 16:03:09 -07:00
Ryan Houdek a8c1a36c12 Docs: Update for release FEX-2504 2025-04-04 14:26:30 -07:00
Ryan Houdek acfb24f871 Merge pull request #4477 from pmatos/incdecstp-tests
Update tests for fincstp and fdecstp
2025-04-04 09:37:56 -07:00
Paulo Matos ecff5aea71 Update tests for fincstp and fdecstp
Followup to FEX-Emu#4475.
Tests were not really testing the interesting instructions.
2025-04-04 17:59:03 +02:00
LC 74cb225ccb Merge pull request #4478 from Sonicadvance1/sha256rnds2_for_reals
OpcodeDispatcher: Implement support for sha256rnds2 using ARM instructions
2025-04-03 11:26:58 -04:00
LC d5db2ccf18 Merge pull request #4475 from Sonicadvance1/x87_stack_bug
x87OptimizationPass: Fixes {Inc,Dec}StackPop
2025-04-03 11:25:53 -04:00
Ryan Houdek cebcf50c65 x87OptimizationPass: Fixes {Inc,Dec}StackPop
On the slow path these were pushing and popping in the wrong direction.
Switch them around to ensure the unittests work.
2025-04-02 14:25:02 -07:00
Ryan Houdek 3a3c9101c7 InstcountCI: Update 2025-04-02 12:51:03 -07:00
Ryan Houdek ec1c7797f4 OpcodeDispatcher: Implement support for sha256rnds2 using ARM instructions
The big one.
2025-04-02 12:51:03 -07:00
Ryan Houdek 313528c34b InstcountCI: Add sha256rnds2 to crypto file. 2025-04-02 12:44:34 -07:00
Ryan Houdek 907fd6b04b unittests/Emitter: Enable crypto tests since vixl supports them now. 2025-04-02 12:43:50 -07:00
Ryan Houdek aa0fc9071e CodeEmitter: Fixes typo in sha256h2
This was accidentally encoding as sha256h.
2025-04-02 12:38:47 -07:00
Tony Wasserka 0bd924eb7e Merge pull request #4460 from Sonicadvance1/dead_code
FEXCore/Frontend: Remove logically dead code
2025-04-02 09:20:41 +02:00
LC 852109c142 Merge pull request #4476 from Sonicadvance1/sha256
IR: Implement support for sha256h{2,}
2025-04-01 21:04:49 -04:00
Ryan Houdek 9b0bb29d78 IR: Implement support for sha256h{2,}
I keep carrying this patch around. Not yet wired up to the instruction
implementation yet, but I don't want to forget about it.
2025-04-01 17:53:48 -07:00
Ryan Houdek c3e71de1d7 FEXCore/Frontend: Remove logically dead code
This code can't get hit.
2025-04-01 16:09:07 -07:00
Ryan Houdek 0256d6820c ASM: Adds a unittest for an x87 stack management bug
This interaction between fxch and fincstp/fdecstp is mind breaking.
2025-04-01 15:03:17 -07:00
Ryan Houdek 1b18bfaff5 Merge pull request #4473 from bylaws/win32-f
Windows: Small fixups
2025-04-01 10:53:28 -07:00
Ryan Houdek 6065e7a62b Merge pull request #4469 from pmatos/InitOutputFD
Initialize OutputFS to -1
2025-04-01 08:53:12 -07:00
Ryan Houdek 8aecdc536c Merge pull request #4471 from alyssarosenzweig/opt/cvtss2si
Optimize float->integer conversions with Feat_FRINTTS
2025-04-01 08:52:50 -07:00
Paulo Matos 86a2e9e655 Initialize OutputFS to -1
Avoids the non-fatal error: "[ERROR] Close closing FEX FD 2"
that happens when the guest program executes a syscall to close fd 2,
and it's not in fex tracked set.
2025-04-01 08:24:09 +02:00
Billy Laws 8f50106187 Context: Fix incorrect ifdef on ARM64EC 2025-03-31 23:57:58 +01:00
Billy Laws 02d3a319f9 OpTables: Disable thunk opcodes on win32
They are of no use here, and are quite frequent in never-taken blocks in Denuvo games
so treating them as invalid avoids wasting some time.
2025-03-31 23:57:58 +01:00
LC d18d0435ae Merge pull request #4472 from Sonicadvance1/remove_warnings_13
Arm64Emitter: Removes warning
2025-03-31 17:09:47 -04:00
Ryan Houdek cdaf1c5262 Arm64Emitter: Removes warning 2025-03-31 13:51:50 -07:00
Alyssa Rosenzweig 4b0e3bff54 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:09:54 -04:00
Alyssa Rosenzweig c4f7b27459 OpcodeDispatcher: accelerate F->I conversions with FRINTTS
this should significantly help perf on supported platforms.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:09:54 -04:00
Alyssa Rosenzweig 3eb8be9953 InstructionCountCI: enable FRINTTS when we use float->int
the relevant target hw (e.g. apple m1) supports this, so let's track with it on.

Secondary_REP and VEX_map1 are duplicated to have frintts and !frintts versions
so we can track the non-frintts path instead, since armv8.5 is still kinda new.
the rest just enable on to minimize combinatorics.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:09:50 -04:00
Alyssa Rosenzweig 1f08f8df0d IR: allow VUShrNI with bitshift=0
encodes to Xtn, we need this to narrow.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Alyssa Rosenzweig 0038a0b19c IR: plumb Vector_FToISized op
this exposes the frint* opcodes in a new ir op

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Alyssa Rosenzweig 166a7c7e53 FEXCore: plumb Feat_FRINTTS
we want these instructions to accelerate conversions.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:01:32 -04:00
Ryan Houdek 8ac296bd6f Merge pull request #4463 from Sonicadvance1/telemetry_remove_indirection
Telemetry: Removes unnecessary indirection
2025-03-31 10:54:52 -07:00
Ryan Houdek e092a38e0f Merge pull request #4464 from Sonicadvance1/cpu_classification
Scripts: Update CPU classification
2025-03-31 10:54:34 -07:00
LC b145e894e4 Merge pull request #4466 from Sonicadvance1/sanitize_programheader_size
ELFParser: Sanity check ELF program headers
2025-03-30 12:23:16 -04:00
LC 63a8b66b28 Merge pull request #4467 from Sonicadvance1/elfcontainer_print_removal
ELFContainer: Removes unused debug print functions
2025-03-30 12:21:39 -04:00
Ryan Houdek a16cc87852 ELFContainer: Removes unused debug print functions
These are completely unused since this ELFContainer is significantly
less utilized than original expectations.

More of this code is dead and can be removed in the future.
2025-03-30 09:01:43 -07:00
Ryan Houdek 1050b60057 ELFParser: Sanity check ELF program headers
Malformed ELF files could parse in a bad offset.
2025-03-30 08:56:14 -07:00
Ryan Houdek 9188e85164 Scripts: Update CPU classification
Hasn't been updated in a while, missing a bunch of CPUs. Noticed since
my Orion-O6 was getting compiled for a Cortex-A57.

Additionally remove the comment about llvm not being able to detect
newer Kryo CPUs since that has been fixed since Clang 12 from MR
https://reviews.llvm.org/D94954 and our minimum spec is Clang 13.
2025-03-29 15:36:17 -07:00
Ryan Houdek f9b369c550 Telemetry: Removes unnecessary indirection
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.

Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).

This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
2025-03-29 15:10:37 -07:00
LC 949b205f42 Merge pull request #4462 from Sonicadvance1/passes_initialize_data
x87StackOptimizationPass: Initialize a couple of arrays
2025-03-29 18:00:47 -04:00
LC 31ee8d8178 Merge pull request #4456 from Sonicadvance1/futimesat_finally
LinuxSyscalls: Emulate futimesat syscalls
2025-03-29 17:59:25 -04:00
LC 3155590e87 Merge pull request #4461 from Sonicadvance1/change_vex_operand_encoding
FEXCore/Frontend: Changes how VEX operand encoding flags are encoded
2025-03-29 17:58:31 -04:00
Ryan Houdek feab0bce4b x87StackOptimizationPass: Initialize a couple of arrays
Just to silence some warnings that think these aren't zero initialized
before using.
2025-03-29 14:02:46 -07:00
Ryan Houdek 0599d80b13 FEXCore/Frontend: Changes how VEX operand encoding flags are encoded
These three options are mutually exclusive with each other and could
potentially result in invalid encodings of the table on accident.

Change over to a 2-bit bitfield to encode if the operand that consumes
the VEX option is none, destination, 1st src, or 2nd src.

This ensures the table can't ever be incorrectly encoded.
2025-03-29 13:51:02 -07:00
Ryan Houdek d58e12c5e4 FEXLinuxTests: Adds futimesat test 2025-03-29 11:06:52 -07:00
Ryan Houdek 68939c5a5c LinuxSyscalls: Emulate futimesat syscalls
Since this syscall doesn't exist, we need to convert it to the
equivalent utimensat like the kernel does internally.

This is fairly trivial but there are some safety nets in place.
2025-03-29 11:06:51 -07:00
Ryan Houdek 25c4fb8508 Merge pull request #4458 from lioncash/deprecated
General: Replace deprecated std::is_trivial/std::is_trivial_v trait usages
2025-03-28 22:36:56 -07:00
Lioncache 88e6c48db7 X86Tables: Replace deprecated std::is_trivial template
This is deprecated in C++26, so we can just use a more specific type trait.
2025-03-29 00:55:15 -04:00
Lioncache 218b0d491a LinuxSyscalls: Replace use of deprecated std::is_trivial template
This type trait is deprecated in C++26, so we can just be more specific.
2025-03-29 00:53:30 -04:00
Lioncache bfba74dab9 IR: Replace use of deprecated std::is_trivial_v template
This is deprecated in C++26
2025-03-29 00:33:39 -04:00
Lioncache 5708846a07 CodeEmitter/Registers: Replace use of deprecated std::is_trivial_v template
std::is_trivial/std::is_trivial_v is deprecated in C++26, so we can just
use a more specific trait and be more explicit about what we want.

Also some of these static_asserts were testing the constraints of the wrong
class, so we can tidy those up as well.
2025-03-29 00:28:14 -04:00
Ryan Houdek fd28783f85 Merge pull request #4457 from lioncash/crypto
CodeEmitter/SVEOps: Add SVE2 crypto operations
2025-03-28 14:55:40 -07:00
Ryan Houdek 0ca34d11ad Merge pull request #4454 from alyssarosenzweig/silly-nop
Fix 66 90 decoding to a nop
2025-03-28 14:50:02 -07:00
Ryan Houdek bf3275ba4a Merge pull request #4451 from pmatos/SingleStepCheck
Enable maxinst to 1 only if singlestep exists and is enabled
2025-03-28 14:45:40 -07:00
Ryan Houdek d8e4e00b2b Merge pull request #4455 from alyssarosenzweig/opt/cvtss2si
InstructionCountCI: add blocks using F->I conversion
2025-03-28 14:45:23 -07:00
Ryan Houdek 8b38a6dd08 Merge pull request #4452 from lioncash/buf
CodeEmitter: Generify data writing
2025-03-28 14:45:02 -07:00
Lioncache 63f72621fc CodeEmitter/SVEOps: Add SVE2 crypto constructive binary operations
Likewise, these are now also available to test against in vixl.
2025-03-28 16:21:21 -04:00
Lioncache 580a8c9c61 CodeEmitter/SVEOps: Add SVE2 crypto destructive binary operations
These are also now available in vixl to run disassembly tests against.
2025-03-28 16:09:41 -04:00
Lioncache 249351cef4 CodeEmitter/SVEOps: Add SVE2 crypto unary operations
Now that these are available in vixl, we finally have something to test against.
2025-03-28 15:59:16 -04:00
Alyssa Rosenzweig 46690ae352 InstructionCountCI: add blocks using F->I conversion
to see the impact of Billy's stuff in a real world context since we don't have
hot blocks for it yet. random blocks I pulled from Control_DX11.exe which I had
handy, with rip relative addressing replaced.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-28 13:14:18 -04:00
Paulo Matos 68b5a90518 Enable maxinst to 1 only if singlestep exists and is enabled 2025-03-28 18:09:52 +01:00
Alyssa Rosenzweig cb54823622 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-28 13:03:34 -04:00
Alyssa Rosenzweig 3a2ca41724 OpcodeDispatcher: handle 66 90 as a NOP
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-28 13:03:34 -04:00
Alyssa Rosenzweig dd4e6d29ca InstructionCountCI: add bytemark neural-net blocks
first doesn't have any low hanging opts, but shows the potential win from post-RA
CSE type opts. whether those are actually a /good/ idea is harder to say, but
it's an idea.

second has some interesting issues here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-28 13:03:31 -04:00
Lioncache 497ee32f59 CodeEmitter: Generify data writing
Lets us tidy up the repeated generation code a little bit
(and also allow for generic data writing, should it ever be needed)
2025-03-28 09:55:35 -04:00
Ryan Houdek d34c287f69 Merge pull request #4450 from lioncash/ilog
Addressing: Remove unnecessary assert in LoadEffectiveAddress()
2025-03-28 06:43:31 -07:00
Ryan Houdek b744ca16e1 Merge pull request #4449 from pmatos/ConfigWarnSA
Remove some warnings from Config.cpp
2025-03-28 06:43:01 -07:00
Paulo Matos 573c262bb2 Remove some warnings from Config.cpp
- unused includes, and
- unused functions.
2025-03-28 14:18:40 +01:00
Lioncache ff2c2f1e1f Addressing: Remove unnecessary assert in LoadEffectiveAddress()
ilog2 already has an assert for this.
2025-03-28 06:08:08 -04:00
Ryan Houdek 8058391955 Merge pull request #4448 from lioncash/uninit
OpcodeDispatcher: Fix potential for uninitialized value use in RCRSmallerOp()
2025-03-28 02:08:20 -07:00
Lioncache 54dbb9f248 OpcodeDispatcher: Fix potential for uninitialized value use in RCRSmallerOp()
If Src isn't a constant, then no value is actually assigned to SrcConst, so this
can result in uninitialized arithmetic being performed
2025-03-28 04:45:40 -04:00
Ryan Houdek 7d1351f402 Merge pull request #4447 from lioncash/syscall
Syscalls: Minor cleanup
2025-03-27 20:41:36 -07:00
Ryan Houdek 21233430aa Merge pull request #4446 from lioncash/cond
OpcodeDispatcher: Remove unnecessary 128-bit check in VPGATHER()
2025-03-27 20:30:49 -07:00
Ryan Houdek bd1d6820b7 Merge pull request #4438 from Sonicadvance1/static_analysis_wars
Static analysis warning fixes
2025-03-27 20:30:37 -07:00
Lioncache 18e84ae2e8 Syscalls: Remove duplicate CLONE_NEWUTS in CloneHandler() 2025-03-27 23:26:56 -04:00
Lioncache f9c29056e6 Syscalls: Put limit check before buffer access in GenerateMap()
Just a trivial fix to plug potential UB
2025-03-27 23:24:54 -04:00
Lioncache 40b1c32008 OpcodeDispatcher: Remove unnecessary 128-bit check in VPGATHER()
This is already guaranteed to be true, since it's checked in the outer if,
so this can just be a regular else statement.
2025-03-27 23:08:59 -04:00
Ryan Houdek b5ed804578 Merge pull request #4444 from lioncash/jitcond
JIT: Simplify SVE 256 operation asserts
2025-03-27 19:46:23 -07:00
Ryan Houdek fbe3a86c4e Merge pull request #4443 from alyssarosenzweig/ra/cleanup
RA: small cleanups
2025-03-27 19:46:14 -07:00
Ryan Houdek 498c86b47d Merge pull request #4442 from lioncash/fmt
Externals: Update fmt to 11.1.4 (from 11.1.0)
2025-03-27 19:46:01 -07:00
Ryan Houdek b8b6f81c44 Merge pull request #4441 from lioncash/move
FEXRootFSFetcher: Move strings in GetDistroInfo()
2025-03-27 19:45:52 -07:00
Ryan Houdek aab9e1b751 Merge pull request #4437 from Sonicadvance1/vdso_parser
VDSOEmulation: Safe fallback if host VDSO can't be parsed
2025-03-27 19:45:39 -07:00
Ryan Houdek ec976f3f75 Merge pull request #4435 from alyssarosenzweig/opt/pall-pf-af
Pair PF/AF when spilling static regs
2025-03-27 19:45:29 -07:00
Alyssa Rosenzweig f169fa5da2 Merge pull request #4445 from lioncash/band
Addressing: Amend binary AND into logical AND in SelectAddressMode()
2025-03-27 18:13:21 -04:00
Lioncache 35ee12e7e9 Addressing: Amend binary AND into logical AND in SelectAddressMode()
Bitwise AND here is a little odd and was likely intended to be a logical AND.
2025-03-27 17:29:08 -04:00
Lioncache 4497ab8844 JIT: Simplify SVE 256 operation asserts
We can make these slightly less verbose.
2025-03-27 16:55:20 -04:00
Alyssa Rosenzweig 9f399f3313 RegisterAllocationPass: rm unused arg to DecodeSRAReg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 16:45:31 -04:00
Alyssa Rosenzweig 3939213336 RegisterAllocationPass: rm useless assertion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 16:45:31 -04:00
Lioncache cc27a0f666 Externals: Update fmt to 11.1.4 (from 11.1.0)
Updates fmt to the latest bugfix release.
2025-03-27 13:52:59 -04:00
Lioncache 1ff2216063 FEXRootFSFetcher: Move strings in GetDistroInfo()
Just a few instances where static analysis reports unnecessary copies.
2025-03-27 13:36:00 -04:00
Ryan Houdek b8165813b4 Merge pull request #4439 from pmatos/Init-SA
Initialize class fields to null/zero
2025-03-27 09:32:18 -07:00
Alyssa Rosenzweig 2215b153db InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 11:07:29 -04:00
Alyssa Rosenzweig cb91d585c3 Arm64Emitter: pair pf/af load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 11:07:00 -04:00
Alyssa Rosenzweig c971d4044f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 11:06:57 -04:00
Ryan Houdek fdb4a078f2 Merge pull request #4440 from pmatos/pylance-warn
Fix pylance warning about possible unbound var
2025-03-27 08:06:39 -07:00
Alyssa Rosenzweig 42ea711850 CoreState: squish and rearrange pf_raw/af_raw
to allow next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 11:04:08 -04:00
Ryan Houdek 96721978c0 Merge pull request #4434 from alyssarosenzweig/bug/pall
Fix x18 register saving
2025-03-27 08:00:58 -07:00
Paulo Matos d64698e4c9 Fix pylance warning about possible unbound var 2025-03-27 15:53:46 +01:00
Paulo Matos 4565f2b689 Initialize class fields to null/zero
Silence a couple of static analyzer warnings.
2025-03-27 15:35:40 +01:00
Alyssa Rosenzweig 5fa7f1d50d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 10:08:12 -04:00
Alyssa Rosenzweig 2cfc42bd6a Arm64Emitter: simplify and fix !preserve_all regs
stop doing weird special cases. just dump all the regs except what aapcs64 says
we don't have to.

this fixes saving x18 across thunks and things. so probably fixes things *cry*

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 10:07:27 -04:00
Ryan Houdek 003de2b659 JIT: static analysis warnings 2025-03-27 04:36:41 -07:00
Ryan Houdek 4ab6c2252d IR: More validation for shifts 2025-03-27 04:36:41 -07:00
Ryan Houdek 8f9d818368 OpcodeDispathcer: static analysis warnings 2025-03-27 04:26:40 -07:00
Ryan Houdek 4ebc307744 IR: Ensure BFE ops can't try to extract element larger than size 2025-03-27 04:26:40 -07:00
Ryan Houdek 0b519b29d9 IR/IRDumper: static analysis warnings 2025-03-27 04:26:40 -07:00
Ryan Houdek 8881e8d96e IR/IREmitter: static analysis warnings 2025-03-27 04:26:40 -07:00
Ryan Houdek 9a99608f68 Passes/ConstProp: static analysis warnings 2025-03-27 04:26:40 -07:00
Ryan Houdek 26c26308db VDSOEmulation: Safe fallback if host VDSO can't be parsed
SHouldn't ever occur.
2025-03-27 03:20:40 -07:00
Tony Wasserka 123c5d809e Merge pull request #4433 from pmatos/JSONValidate
Improve JSON file validation and error reporting
2025-03-27 09:57:37 +01:00
Alyssa Rosenzweig 7c42c7798c Arm64Emitter: fix a bunch of preserve_all
* x18 wasn't getting spilled even though it was supposed to be.
* arm64ec preserve_all definitions were all messed up, specialize these to fix.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-26 16:47:39 -04:00
Paulo Matos 0f98daf1d9 Improve JSON file validation and error reporting
Turn invalid JSON files into fatal errors.
2025-03-26 16:17:20 +01:00
LC c3de7c63b4 Merge pull request #4432 from Sonicadvance1/fix_typo2
JIT: Fix typo in VSha256U1
2025-03-25 22:02:55 -04:00
Ryan Houdek ab9123a427 JIT: Fix typo in VSha256U1 2025-03-25 18:19:37 -07:00
Ryan Houdek e504a8c979 Merge pull request #4428 from alyssarosenzweig/opt/vbsl-tie
Tie VBSL source
2025-03-25 16:09:17 -07:00
Ryan Houdek f6dd87a3a1 Merge pull request #4430 from alyssarosenzweig/opt/bfi-zext
JIT: eliminate a zext in bfi
2025-03-25 16:08:33 -07:00
Alyssa Rosenzweig 89c530054d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 17:37:29 -04:00
Alyssa Rosenzweig f87edbe1cb JIT: eliminate a zext in bfi
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 17:36:59 -04:00
Alyssa Rosenzweig 4aa477de3e InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:23:58 -04:00
Alyssa Rosenzweig 694e674fe6 IR: tie VExtr
needed for sve-256 move reduction.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:22:55 -04:00
Alyssa Rosenzweig 7ebc0f32b8 IR: tie VInsGPR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:16:25 -04:00
Alyssa Rosenzweig 68cacc2fc3 IR: tie VFMin/VFMax
this was missed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:16:25 -04:00
Alyssa Rosenzweig 43eb597044 IR: tie VBSL source
this was missed before, noticed while experimenting with round robin RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 14:14:21 -04:00
Tony Wasserka 7efd827e78 Merge pull request #4378 from bylaws/volmd
Implement PE volatile metadata support
2025-03-25 10:23:53 +01:00
Ryan Houdek 2f6f8b93e9 Merge pull request #4424 from Sonicadvance1/converted_options
Convert config options once
2025-03-24 19:49:32 -07:00
Ryan Houdek d48413cbec Merge pull request #4426 from alyssarosenzweig/opt/shld-improvements
Drop some masking in shld
2025-03-24 19:24:51 -07:00
Ryan Houdek 6bafca688b Merge pull request #4427 from alyssarosenzweig/opt/cmpxchg-trivial
OpcodeDispatcher: drop useless mask in trivial cmpxchg
2025-03-24 19:24:39 -07:00
Ryan Houdek 53356f1aa7 Merge pull request #4422 from Sonicadvance1/wine_support_sleep
WINE: Support sleeping a process
2025-03-24 19:19:38 -07:00
Billy Laws 51281f6a3a ARM64EC: Load volatile metadata 2025-03-24 22:01:49 +00:00
Billy Laws 7a5e08c5ab FEXCore: Use TSO range information when emitting IR 2025-03-24 22:01:49 +00:00
Billy Laws 642903a7bf FEXCore: Support tracking TSO range information 2025-03-24 22:01:49 +00:00
Billy Laws f51fd6c78d Move IntervalList to FEXCore 2025-03-24 22:01:49 +00:00
Billy Laws 03cf15a9e1 Windows: Add IntervalList batch Insert and Contains methods 2025-03-24 22:01:49 +00:00
Alyssa Rosenzweig fe19c04f6a InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-24 16:17:39 -04:00
Alyssa Rosenzweig b86cbda03d OpcodeDispatcher: drop useless mask in trivial cmpxchg
hit by upcoming opt pass, but we don't want to depend on that pass for that.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-24 16:17:04 -04:00
Alyssa Rosenzweig b1af6e23cf InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-24 16:07:43 -04:00
Alyssa Rosenzweig e26b9b12fa OpcodeDispatcher: drop useless zext in shld
this gets deleted by an upcoming opt pass, but we shouldn't be depending on
the opt pass for it!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-24 16:00:24 -04:00
Ryan Houdek 9bf47b3f23 Convert config options once
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.

Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
2025-03-22 17:32:10 -07:00
LC 2887416b2e Merge pull request #4425 from Sonicadvance1/recalibrate_your_senses
FEXCore: Remove CPUID option for SHA
2025-03-22 18:59:56 -04:00
Ryan Houdek 9a92f6f743 FEXCore: Remove CPUID option for SHA
Instead use the HostFeatures option that exists which is controlled by
the {enable,disable}crypto option.
2025-03-21 19:43:38 -07:00
LC 4929480719 Merge pull request #4423 from Sonicadvance1/get_out_of_here_xbyak
FEX: Remove xbyak dependency
2025-03-21 22:04:22 -04:00
Ryan Houdek c8437d2303 WINE: Support sleeping a process
Support using an app profile for putting a process to sleep at the
start. Useful for getting a debugger attached early.
2025-03-21 12:48:38 -07:00
Ryan Houdek ef25ae4d3b FEX: Remove xbyak dependency
Get out of here xbyak.
2025-03-20 19:37:42 -07:00
Ryan Houdek 9399790c11 unittests/ASM: Fix tests expecting a working stack
HostRunner and FEX may or may not set stack. Fix the cases that assumed
it was correctly setup.
2025-03-20 16:54:56 -07:00
Ryan Houdek ab7f6484cf Merge pull request #4414 from pmatos/AddrMode-DeusEx
Fix address modes calculation on 32bit guests
2025-03-20 13:21:53 -07:00
Paulo Matos 2d8c5da379 instcountci: Abstract AddressMode into its own header 2025-03-20 11:41:31 +01:00
Paulo Matos b18f148575 Abstract AddressMode into its own header
Also refactor usage of utility functions into x87 Stack Optimization Pass.
2025-03-20 11:41:26 +01:00
Ryan Houdek 6426428718 Merge pull request #4419 from neobrain/fix_vma_order
SMCTracking: Fix order of VMAs tracked per MappedResource
2025-03-19 12:29:23 -07:00
Ryan Houdek af2ee426f5 Merge pull request #4418 from neobrain/refactor_cppoptparse
CMake: Propagate cpp-optparse include directories automatically
2025-03-19 12:29:04 -07:00
Tony Wasserka 6c4e9ff42d CMake: Propagate cpp-optparse include directories automatically 2025-03-19 12:19:03 +01:00
Tony Wasserka 62a37a7d70 SMCTracking: Fix order of VMAs tracked per MappedResource
Previously, new VMA entries were always prepended to the list of the
associated MappedResource. This usually made FirstVMA erroneously point to
the *highest* VMA instead of the lowest.
2025-03-19 12:09:21 +01:00
Ryan Houdek ef83addc74 Merge pull request #4417 from alyssarosenzweig/bug/bt-flags
OpcodeDispatcher: fix BT flags
2025-03-18 23:00:55 -07:00
Alyssa Rosenzweig fb308f5947 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-18 15:26:47 -04:00
Alyssa Rosenzweig cedb93c11c unittests: add test for BT preserving Z
we were clobbering. this test fails on upstream.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-18 15:26:47 -04:00
Alyssa Rosenzweig 61d2a09827 OpcodeDispatcher: fix BT flags
ZF needs to be preserved.

the new code is the same instr count on flagm although probably an extra uop.
the inst count regression is on flagm, but we can't tolerate broken behaviour.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-18 15:26:47 -04:00
Paulo Matos 9c6423f37a instcountci: Do not attempt direct encoding of 32bit addressing modes 2025-03-18 18:34:09 +01:00
Paulo Matos e18a661b50 Refactor OP_STORESTACKMEM case in x87 Stack Opt Pass 2025-03-18 18:34:06 +01:00
Paulo Matos 16e5777816 Do not attempt direct encoding of 32bit addressing modes
Fixes #4393
2025-03-18 15:54:54 +01:00
Paulo Matos 64020e8828 Simplify by merging FSTF64 with FST 2025-03-18 15:54:54 +01:00
Ryan Houdek 30798556fc Merge pull request #4415 from alyssarosenzweig/opt/avx-zero
ConstProp: optimize StoreContext(128-bit zero)
2025-03-17 14:02:06 -07:00
Alyssa Rosenzweig cd3518d99d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-17 16:49:05 -04:00
Alyssa Rosenzweig bc87c3d494 ConstProp: optimize StoreContext(128-bit zero)
Turn this into stp to save a move in a /ton/ of AVX-128 code. It's pretty
annoying to optimize this at the dispatcher level because of the SRA cache, so
doing it in the ConstProp is a nice compromise solution.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-17 16:49:05 -04:00
Ryan Houdek 726228c418 Merge pull request #4409 from Sonicadvance1/fix_memmove_warnings
IoctlEmulation: Stop using std::pair and memmove
2025-03-17 13:08:29 -07:00
Ryan Houdek 929b111648 Merge pull request #4405 from Sonicadvance1/static_analysis_again
Various: More static analysis warnings cleanup
2025-03-17 13:08:11 -07:00
Alyssa Rosenzweig 1700a73382 Merge pull request #4406 from Sonicadvance1/minor_x87_opt
X87: Minor optimization in how GPRs are moved in to vector registers
2025-03-17 10:31:46 -04:00
Tony Wasserka 0f9d791911 Merge pull request #4412 from Sonicadvance1/switch_actions
actions: Update to step-security change-files action
2025-03-17 14:40:45 +01:00
Tony Wasserka 61acc76be7 Merge pull request #4408 from Sonicadvance1/fix_thunks_llvm_20
LibraryForwarding/gen: Fixes compiling with LLVM 20
2025-03-17 14:37:26 +01:00
Ryan Houdek 0f8ba5bf32 actions: Update to step-security change-files action
This repo is provided by them with the commit prior to when it was
compromised. LLVM has switched to the same repo for now.

We probably want to find some better solution to this at some point.
2025-03-15 16:15:59 -07:00
Ryan Houdek 2cb8a96f0e IoctlEmulation: Stop using std::pair and memmove
Fixes a new warning in clang-20 about using memmove on non-trivially
copyable types.
2025-03-14 15:28:53 -07:00
Ryan Houdek 1102122639 ThunkGen: Fixes compiling with LLVM 20
API has changed slightly.

Fixes #4407
2025-03-14 15:14:30 -07:00
Ryan Houdek 66e026a9e8 InstcountCI: Update 2025-03-12 22:23:03 -07:00
Ryan Houdek b901b42417 OpcodeDispatcher/X87: Optimize GPR moves using new IR operation
This saves one instruction per FILD and FRSTOR
2025-03-12 22:22:03 -07:00
Ryan Houdek 9f2f10a65f IR: Add new 128-bit vector move operation
This allows us to slightly optimize some x87 behaviour.
2025-03-12 22:21:28 -07:00
Ryan Houdek b0b41d00ee Various: More static analysis warnings cleanup
NFC
2025-03-12 17:27:41 -07:00
Ryan Houdek 2d56f5eba0 Merge pull request #3487 from neobrain/feature_libfwd_vulkan32
Library Forwarding: Add experimental support for 32-bit Vulkan
2025-03-12 09:50:35 -07:00
Tony Wasserka 8ffc6fbc6b LibraryForwarding/vulkan: Disable 32-bit guest library
This currently doesn't export enough symbols to be viable for practical use.
Building the host library only serves to ensures the relevant features
continue to work however.
2025-03-12 17:35:07 +01:00
Tony Wasserka 0ffd94dadb LibraryForwarding/vulkan: Enable more entry points 2025-03-12 17:35:07 +01:00
Tony Wasserka 585320093a LibraryForwarding: Add custom repacking tests for Vulkan-like scenarios 2025-03-12 17:35:07 +01:00
Tony Wasserka bdd351a42c LibraryForwarding/vulkan: Apply exit repacking for arrays of types that use custom repacking 2025-03-12 17:30:43 +01:00
Tony Wasserka feaee702e9 LibraryForwarding/vulkan: Reload function pointers on device change when needed
This only needs to be done for functions that have a custom host-side
implementation and that don't take a VkDevice argument.
2025-03-12 17:30:43 +01:00
Tony Wasserka d89fc84fcf LibraryForwarding/vulkan: Minor cleanups 2025-03-12 17:30:43 +01:00
Tony Wasserka 688cd1a4bb LibraryForwarding/vulkan: Add support for VK_EXT_descriptor_buffer on 32-bit 2025-03-12 17:30:43 +01:00
Tony Wasserka c056875a00 LibraryForwarding/gen: Allow annotating non-pointer members as custom_repack
This is useful in particular for union members. A guest_layout specialization
must be provided manually for the type of the annotated member.
2025-03-12 17:30:43 +01:00
Tony Wasserka 054118139f LibraryForwarding/vulkan: Implement vkCmdSetVertexInputEXT on 32-bit 2025-03-12 17:30:43 +01:00
Tony Wasserka 447148d95a LibraryForwarding/vulkan: Manually add pointers only referenced through nested pointers 2025-03-12 17:30:43 +01:00
Tony Wasserka f6b1e42a0b LibraryForwarding/vulkan: Enable X11 WSI functions 2025-03-12 17:30:43 +01:00
Tony Wasserka 90e74e5572 LibraryForwarding/vulkan: Disable recently added 1.4 functions on 32-bit guests 2025-03-12 17:30:43 +01:00
Tony Wasserka e4fc5fe5be LibraryForwarding/vulkan: Extend 32-bit support 2025-03-12 17:30:43 +01:00
Tony Wasserka 809b2c6115 LibraryForwarding/vulkan: Add 32-bit support 2025-03-12 17:30:43 +01:00
Tony Wasserka 659538ef4b LibraryForwarding/vulkan: Enable 32-bit build 2025-03-12 17:24:08 +01:00
Tony Wasserka 4b677f4b42 LibraryForwarding: Add array specializations for guest_layout/host_layout
This allows automatic repacking of structs with array members.
2025-03-12 17:24:08 +01:00
Tony Wasserka 87699ea5a0 LibraryForwarding/gen: Skip type compatibility checking when emit_layout_wrappers is used 2025-03-12 17:24:08 +01:00
Tony Wasserka a2ae113ee9 LibraryForwarding/gen: Make second parameter to fex_apply_custom_repacking_exit const 2025-03-12 17:24:08 +01:00
Tony Wasserka 901e2c75d4 FEXLinuxTests: Drop unneeded code 2025-03-12 17:24:08 +01:00
Ryan Houdek 871d140b7c Merge pull request #4404 from neobrain/fix_libfwd_vulkan_missing_decl
LibraryForwarding/vulkan: Add missing vkCmdPushDescriptorSetWithTemplate declaration
2025-03-12 09:18:12 -07:00
Tony Wasserka 2f6ae1ad02 LibraryForwarding/vulkan: Add missing vkCmdPushDescriptorSetWithTemplate declaration 2025-03-12 17:09:11 +01:00
Tony Wasserka 806e98925c Merge pull request #4402 from Sonicadvance1/new_vulkan_headers
LibraryForwarding/vulkan: Update to 1.4.310
2025-03-12 16:26:50 +01:00
Ryan Houdek c7dbd2fac2 Thunks/Vulkan: Update to 1.4.310
Adds support for Vulkan 1.4 and a couple of NVIDIA extensions.

Last update was nearly six months ago in #4090. Once again following the
information to extract definitions from #2076
2025-03-12 08:13:43 -07:00
Ryan Houdek 1b7729efed Scripts/DefinitionExtract: Update to use isystem include
Fixes compiles on newer compiles
2025-03-12 07:45:34 -07:00
Ryan Houdek 6d1d5aeffb External: Update Vulkan-Headers to v1.4.310 2025-03-12 07:45:34 -07:00
Ryan Houdek 9ae04b5771 Merge pull request #4403 from neobrain/refactor_script_format
Scripts/DefinitionExtract: Make output consistent with clang-format
2025-03-12 07:43:08 -07:00
Tony Wasserka fc61e3b1b5 Scripts/DefinitionExtract: Make output consistent with clang-format 2025-03-12 10:15:28 +01:00
Ryan Houdek fa1d9910e4 Merge pull request #4401 from OFFTKP/fcomi
Add tests for F(U)COMI(P)
2025-03-11 17:35:06 -07:00
offtkp d425873eed Use a normal QNaN 2025-03-12 02:27:06 +02:00
offtkp d89c54dd95 That was an sNaN 2025-03-12 01:00:50 +02:00
offtkp 5e023c55bd Fix endianness 2025-03-12 00:46:14 +02:00
offtkp 2487172df3 Initial test 2025-03-12 00:30:16 +02:00
LC bdd078df17 Merge pull request #4400 from Sonicadvance1/struct_verifier_isystem
StructVerifier: Use isystem for system header includes
2025-03-10 21:34:28 -04:00
Ryan Houdek db75335ad0 StructVerifier: Use isystem for system header includes
Fixes some errors with newer compilers.
2025-03-10 13:31:08 -07:00
Ryan Houdek 531ab5b5b1 Merge pull request #4399 from OFFTKP/br
Remove some brackets from expected results in tests
2025-03-10 11:40:31 -07:00
offtkp 0601a863f9 Don't use brackets for GPRs in x87 tests 2025-03-10 18:34:40 +02:00
LC 10d34c9564 Merge pull request #4397 from Sonicadvance1/optimize_cas
JIT: Optimize CAS
2025-03-08 12:06:49 -05:00
Ryan Houdek a37d6a3841 JIT: Optimize CAS
Hey kid, want to see a sick trick?

Finally optimal codegen for 64-bit cmpxchg.
2025-03-07 14:24:45 -08:00
LC d5234a43da Merge pull request #4396 from Sonicadvance1/remove_old_comment
AVX128: Remove old comment from VEXTRACT{F,I}128
2025-03-07 17:20:23 -05:00
LC b7733540c1 Merge pull request #4395 from Sonicadvance1/missing_spdx
Various: Adds missing SPDX file headers
2025-03-07 15:32:47 -05:00
Ryan Houdek 2ae02ded74 AVX128: Remove old comment from VEXTRACT{F,I}128
This comment isn't relevant anymore as unused half of ymm loads will get
DCE'd
2025-03-07 12:00:35 -08:00
LC 2e57ad644d Merge pull request #4394 from Sonicadvance1/nitty_nit
unittests/ASM: Move test to correct folder
2025-03-07 14:38:14 -05:00
Ryan Houdek 031afbfb18 Various: Adds missing SPDX file headers
NFC
2025-03-07 11:34:31 -08:00
Ryan Houdek 377ce2e2f6 unittests/ASM: Move test to correct folder
This is an H0F3A operation not H0F38
2025-03-07 11:17:13 -08:00
LC a2fd3c077d Merge pull request #4390 from Sonicadvance1/fix_inline_softfloat
Softfloat: Define INLINE
2025-03-07 12:32:01 -05:00
Ryan Houdek 0f5aff73ab Softfloat: Define INLINE
This is embarassing. We were throwing away performance by failing to
use the softfloat library's inline helpers.

Turns out we needed to define `INLINE` to something in order for them to
work.

Feels bad.
2025-03-07 01:56:13 -08:00
LC f0b208e692 Merge pull request #4391 from Sonicadvance1/optimize_vpalignr
AVX128: Optimize vpalignr
2025-03-07 02:11:18 -05:00
LC 7fcb5fc590 Merge pull request #4392 from Sonicadvance1/fix_scalar_recip
JIT: Fixe scalar reciprocal when AFP is supported
2025-03-07 02:10:20 -05:00
Ryan Houdek b76f819759 InstcountCI: Update 2025-03-06 18:48:43 -08:00
Ryan Houdek 4b36d4f1ea JIT: Fixe scalar reciprocal when AFP is supported
Somehow I had completely missed this and recent reciprocal tests have
exposed it as a problem. When AFP is supported but not RPRES then we
were hitting this code path.

We were failing to insert in to the destination correctly, which because
the reciprocal is calculated using fdiv using a synthesized constant,
this would just zero the remaining portion of the register.
2025-03-06 18:21:26 -08:00
Ryan Houdek f4d0c6c807 InstcountCI: Update 2025-03-06 17:13:30 -08:00
Ryan Houdek 79ab76b42e AVX128: Optimize vpalignr
When the shift size is exactly 16bytes, then it turns in to a move.

If the shift size is above 16-bytes then synthesize the zero register in
the OpcodeDispatcher, so the backend doesn't synthesize and not cache.
2025-03-06 17:10:22 -08:00
Ryan Houdek 8c3ac61f8d InstcountCI: Move FMA4 to its own file 2025-03-06 17:04:01 -08:00
LC a8bc20f76b Merge pull request #4385 from Sonicadvance1/values_savings
Config: Stop using config values with list when unnecessary
2025-03-06 18:18:51 -05:00
LC 53a55baf23 Merge pull request #4389 from Sonicadvance1/df_on_signal
SignalDelegator: Clear DF and RF on signal
2025-03-06 09:22:30 -05:00
Ryan Houdek 63d6800cf3 unittests/FEXLinuxTests: Ensure that DF is reset inside signal handler 2025-03-05 19:10:54 -08:00
Ryan Houdek 44492b4828 SignalDelegator: Clear DF and RF on signal
The Linux kernel clears these flags on signal, DF is particularly
dangerous because it would break ABI if a signal happened to occur in
the middle of a memory operation that changed the direction of copy.

Because of how frequently wine uses signals, this is actually fairly
likely to occur inside of a memcpy/memset function.

Shout out to BlinkDagger on Discord who found that we forgot to do this.
2025-03-05 18:52:23 -08:00
LC 8ed9f8aef5 Merge pull request #4387 from Sonicadvance1/analysis_warnings
IR: Remove some static analysis warnings
2025-03-05 21:01:46 -05:00
Ryan Houdek 711021e24e Merge pull request #4388 from bylaws/new-mingw
Windows: Support newer mingw toolchains
2025-03-05 15:28:18 -08:00
LC 55dea335fe Merge pull request #4386 from Sonicadvance1/optimize_masked_contiguous
JIT: Optimize SVE offset VL loadstores
2025-03-05 17:18:12 -05:00
Billy Laws 7097532ddf Windows: Support newer mingw toolchains 2025-03-05 21:24:30 +00:00
Ryan Houdek 66841ce35c IR: Remove some static analysis warnings
NFC
2025-03-05 12:30:27 -08:00
Ryan Houdek 3d814cb7c1 InstcountCI: Update 2025-03-05 12:25:33 -08:00
Ryan Houdek 07afdca58b V{Load,Store}VectorMasked: Small offset support with ASIMD 2025-03-05 12:25:33 -08:00
Ryan Houdek 93ed346fd2 JIT: Optimize SVE offset VL loadstores
This was only wired up for 256-bit SVE and wasn't ever hit for 128-bit
SVE. Ensure it works with 128-bit SVE, so mulvl needs to know when
128-bit is used. Then wire it up for vmaskmovps/pd. This saves one
instruction per operation.

Fixes #3791.
2025-03-05 12:01:50 -08:00
Ryan Houdek c9eee9bf7f Config: Stop using config values with list when unnecessary
With the previous fixes in place, we can now stop burning a fextl::list
in every single config option. This list is only required for strarray
options so reserve it for those entirely.

We also don't need to save the config option enum for each, so these
actually go from ~32 bytes per object down to their base type for most
everything.
2025-03-05 11:17:09 -08:00
663 changed files with 74801 additions and 196904 deletions

No files matched your search

+2 -2
View File
@@ -32,7 +32,7 @@ AttributeMacros:
BinPackArguments: true
BinPackParameters: true
BitFieldColonSpacing: Both
BreakAfterAttributes: Always # clang 16 required
BreakAfterAttributes: Leave
BreakBeforeBraces: Attach
BreakBeforeBinaryOperators: None
BreakBeforeInlineASMColon: OnlyMultiline # clang 16 required
@@ -60,7 +60,7 @@ IndentRequires: false
IndentWidth: 2
InsertBraces: true
KeepEmptyLinesAtTheStartOfBlocks: true
LambdaBodyIndentation: OuterScope
LambdaBodyIndentation: Signature
LineEnding: LF # clang 16 required
MaxEmptyLinesToKeep: 2
NamespaceIndentation: Inner
-6
View File
@@ -1,10 +1,4 @@
# This file is used to ignore files and directories from clang-format
# Ignore all files in the External directory
External/*
# SoftFloat-3e code doesn't belong to us
FEXCore/Source/Common/SoftFloat-3e/*
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
+4
View File
@@ -16,3 +16,7 @@
# Reformat of CodeEmitter inl files
8760c593ece92d7e9fa94c40da0368fd367c9cad
# Whole-tree reformat with clang-format-19
5267cde60e7642852d18f20ae8568643bb5293d5
+1 -1
View File
@@ -78,7 +78,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+6 -12
View File
@@ -28,7 +28,7 @@ jobs:
- name: Get changed files
id: changed-files
uses: tj-actions/changed-files@v39
uses: step-security/changed-files@3dbe17c78367e7d60f00d78ae6781a35be47b4a1 # v45.0.1
with:
separator: ","
skip_initial_fetch: true
@@ -40,11 +40,8 @@ jobs:
echo "Formatting files:"
echo "$CHANGED_FILES"
- name: Check for correct clang-format version
run: clang-format --version | grep -qF '16.0.6'
- name: Check git-clang-format-16 exists
run: which git-clang-format-16
- name: Check git-clang-format-19 exists
run: which git-clang-format-19
- name: Setup Python env
uses: actions/setup-python@v4
@@ -58,19 +55,16 @@ jobs:
- name: Run code formatter
env:
CLANG_FORMAT_PATH: 'git-clang-format-16'
CLANG_FORMAT_PATH: 'git-clang-format-19'
GITHUB_PR_NUMBER: ${{ github.event.pull_request.number }}
START_REV: ${{ github.event.pull_request.base.sha }}
END_REV: ${{ github.event.pull_request.head.sha }}
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
# TODO(pmatos): Once we adopt v18, we should be able
# to take advantage of the new --diff_from_common_commit option
# explicitly in code-format-helper.py and not have to diff starting at
# the merge base.
# Using --diff_from_common_commit option available in clang-format-19
run: |
python ./External/code-format-helper/code-format-helper.py \
--repo "FEX-emu/FEX" \
--issue-number $GITHUB_PR_NUMBER \
--start-rev $(git merge-base $START_REV $END_REV) \
--start-rev $START_REV \
--end-rev $END_REV \
--changed-files "$CHANGED_FILES"
+88
View File
@@ -0,0 +1,88 @@
name: Wine DLL artifacts
on:
push:
branches:
- main
env:
BUILD_TYPE: Release
jobs:
wine_dll_artifacts:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, mingw]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Add MingGW to PATH
run: echo "$HOME/llvm-mingw/build/bin/" >> $GITHUB_PATH
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean install directory
run: |
rm -Rf ${{runner.workspace}}/build_install
mkdir ${{runner.workspace}}/build_install
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build_arm64ec
rm -Rf ${{runner.workspace}}/build_wow64
- name: Create Build Environment arm64ec
run: |
cmake -E make_directory ${{runner.workspace}}/build_arm64ec
cmake -E make_directory ${{runner.workspace}}/build_wow64
- name: Configure CMake arm64ec
shell: bash
working-directory: ${{runner.workspace}}/build_arm64ec
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Configure CMake wow64
shell: bash
working-directory: ${{runner.workspace}}/build_wow64
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Build arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Build wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
overwrite: true
name: wine_dll_artifacts
path: ${{runner.workspace}}/build_install/usr/lib/wine/aarch64-windows/lib*.dll
retention-days: 60
compression-level: 9
-4
View File
@@ -5,10 +5,6 @@
[submodule "External/cpp-optparse"]
path = Source/Common/cpp-optparse
url = https://github.com/Sonicadvance1/cpp-optparse
[submodule "External/xbyak"]
shallow = true
path = External/xbyak
url = https://github.com/herumi/xbyak.git
[submodule "External/fex-posixtest-bins"]
shallow = true
path = External/fex-posixtest-bins
+39 -15
View File
@@ -13,6 +13,7 @@ option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
set(USE_LINKER "" CACHE STRING "Allow overriding the linker path directly")
option(ENABLE_UBSAN "Enables Clang UBSAN" FALSE)
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_COVERAGE "Enables Coverage" FALSE)
@@ -34,10 +35,19 @@ set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use fo
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_DEV_ROOTFS "/" CACHE FILEPATH "Path to the sysroot used for cross-compiling for i686 and x86_64")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
set (DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
set (HOSTLIBS_DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
if (NOT DATA_DIRECTORY)
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu")
endif()
include(GNUInstallDirs)
if (NOT HOSTLIBS_DATA_DIRECTORY)
set(HOSTLIBS_DATA_DIRECTORY "${CMAKE_INSTALL_FULL_LIBDIR}/fex-emu")
endif()
string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
@@ -93,7 +103,7 @@ endif()
# uninstall target
if(NOT TARGET uninstall)
configure_file(
"${CMAKE_CURRENT_SOURCE_DIR}/CMakeFiles/cmake_uninstall.cmake.in"
"${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/cmake_uninstall.cmake.in"
"${CMAKE_CURRENT_BINARY_DIR}/CMakeFiles/cmake_uninstall.cmake"
IMMEDIATE @ONLY)
@@ -228,6 +238,19 @@ if (NOT ENABLE_OFFLINE_TELEMETRY)
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
if (ENABLE_UBSAN)
# See https://github.com/FEX-Emu/FEX/pull/4494#issuecomment-2800608944
# and related discussion for the use of -fno-sanitize=alignment -fno-sanitize=function
# with UBSAN.
# alignment: we don't follow a strict alignment policy, for example IR uses packed structs
# that are regularly access unaligned.
# function: syscalls cast function pointers to void (*)(unsigned long...), causing warnings
# related to this access.
add_definitions(-DENABLE_UBSAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=undefined -fno-sanitize=alignment -fno-sanitize=function -fno-sanitize-recover=undefined)
link_libraries(-fno-omit-frame-pointer -fsanitize=undefined -fno-sanitize=alignment -fno-sanitize=function -fno-sanitize-recover=undefined)
endif()
if (ENABLE_ASAN)
add_definitions(-DENABLE_ASAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
@@ -384,15 +407,18 @@ if (TUNE_CPU STREQUAL "native")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=native")
endif()
else()
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo" "${CMAKE_CXX_COMPILER_VERSION}"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/NeedDisabledSVE.py"
RESULT_VARIABLE NEEDS_SVE_DISABLED)
if (NEEDS_SVE_DISABLED)
message(STATUS "Platform has bugged SVE. Disabling")
set(AARCH64_CPU "cortex-a78")
endif()
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${AARCH64_CPU}")
@@ -404,7 +430,7 @@ if (TUNE_CPU STREQUAL "native")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=native")
endif()
endif()
else()
elseif (NOT TUNE_CPU STREQUAL "none")
check_cxx_compiler_flag("-mcpu=${TUNE_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${TUNE_CPU}")
@@ -423,10 +449,6 @@ endif()
add_compile_options(-Wall)
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/include/Config.h.in
${CMAKE_BINARY_DIR}/generated/ConfigDefines.h)
include(CTest)
if (BUILD_TESTS)
message(STATUS "Unit tests are enabled")
@@ -444,6 +466,8 @@ if (BUILD_TESTS)
set(TEST_JOB_FLAG "-j${TEST_JOB_COUNT}")
endif()
add_subdirectory(External/SoftFloat-3e/)
add_subdirectory(External/cephes/)
add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
@@ -607,12 +631,12 @@ set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.com>")
set (CPACK_PACKAGE_VERSION_MAJOR "${FEX_VERSION_MAJOR}")
set (CPACK_PACKAGE_VERSION_MINOR "${FEX_VERSION_MINOR}")
set (CPACK_PACKAGE_VERSION_PATCH "${FEX_VERSION_PATCH}")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/CPack/Description.txt")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/Description.txt")
# Debian defines
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libc6, libstdc++6, libepoxy0, libsdl2-2.0-0, libegl1, libx11-6, squashfuse")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA
"${CMAKE_CURRENT_SOURCE_DIR}/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/CPack/triggers")
"${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/triggers")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# binfmt_misc conflicts with qemu-user-static
# We also only install binfmt_misc on aarch64 hosts
-3
View File
@@ -1,3 +0,0 @@
x86 and x86-64 Linux emulator
FEX is very much work in progress, so expect things to change.
File diff suppressed because it is too large. Load diff
+17 -20
View File
@@ -3,6 +3,7 @@
#include <cstddef>
#include <cstdint>
#include <cstring>
#include <type_traits>
namespace ARMEmitter {
class Buffer {
@@ -21,42 +22,38 @@ public:
Size = BaseSize;
}
template<typename T>
requires (std::is_trivially_copyable_v<T>)
void dcn(const T& Data) {
std::memcpy(CurrentOffset, &Data, sizeof(Data));
CurrentOffset += sizeof(Data);
}
void dc8(uint8_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
dcn(Data);
}
void dc16(uint16_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
dcn(Data);
}
void dc32(uint32_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
dcn(Data);
}
void dc64(uint64_t Data) {
dcn(Data);
}
void dc64(uint64_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void EmitString(const char* String) {
const auto StringLength = strlen(String);
memcpy(CurrentOffset, String, StringLength);
CurrentOffset += StringLength;
}
void Align() {
// Align the buffer to instruction size
auto CurrentAlignment = reinterpret_cast<uint64_t>(CurrentOffset) & 0b11;
void Align(size_t Size = 4) {
// Align the buffer to provided size.
auto CurrentAlignment = reinterpret_cast<uint64_t>(CurrentOffset) & (Size - 1);
if (!CurrentAlignment) {
return;
}
CurrentOffset += 4 - CurrentAlignment;
CurrentOffset += Size - CurrentAlignment;
}
template<typename T>
+13
View File
@@ -86,6 +86,14 @@ constexpr size_t SubRegSizeInBits(SubRegSize size) {
return size_t {8} << FEXCore::ToUnderlying(size);
}
// Many floating point operations constrain their element sizes to the
// main three float sizes half, single, and double precision. This just
// combines all the checks together for brevity.
[[nodiscard]]
constexpr bool IsStandardFloatSize(SubRegSize size) {
return size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit;
}
/* This `ScalarRegSize` enum is used for most scalar float
* operations.
*
@@ -355,6 +363,7 @@ enum class SystemRegister : uint32_t {
TPIDRRO_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b011>,
CNTFRQ_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b000>,
CNTVCT_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b010>,
CNTVCTSS_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b110>,
};
template<uint32_t op1, uint32_t CRm, uint32_t op2>
@@ -572,6 +581,10 @@ enum class Rotation : uint32_t {
template<typename T>
concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegister>;
// Concept for contraining some instructions to accept only a QRegister or DRegister.
template<typename T>
concept IsQOrDRegister = std::is_same_v<T, QRegister> || std::is_same_v<T, DRegister>;
// Whether or not a given set of vector registers are sequential
// in increasing order as far as the register file is concerned (modulo its size)
//
+414 -12
View File
@@ -416,8 +416,7 @@ public:
0;
ASIMDSTLD<size, true, 1>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld1r(T rt, Register rn) {
constexpr uint32_t Op = 0b0000'1101'000 << 21;
constexpr uint32_t Opcode = 0b110;
@@ -435,8 +434,7 @@ public:
0;
ASIMDSTLD<size, true, 2>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld2r(T rt, T rt2, Register rn) {
LOGMAN_THROW_A_FMT(AreVectorsSequential(rt, rt2), "rt and rt2 must be sequential");
constexpr uint32_t Op = 0b0000'1101'000 << 21;
@@ -455,8 +453,7 @@ public:
0;
ASIMDSTLD<size, true, 3>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld3r(T rt, T rt2, T rt3, Register rn) {
LOGMAN_THROW_A_FMT(AreVectorsSequential(rt, rt2, rt3), "rt, rt2, and rt3 must be sequential");
constexpr uint32_t Op = 0b0000'1101'000 << 21;
@@ -475,8 +472,7 @@ public:
0;
ASIMDSTLD<size, true, 4>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld4r(T rt, T rt2, T rt3, T rt4, Register rn) {
LOGMAN_THROW_A_FMT(AreVectorsSequential(rt, rt2, rt3, rt4), "rt, rt2, rt3, and rt4 must be sequential");
constexpr uint32_t Op = 0b0000'1101'000 << 21;
@@ -1652,8 +1648,8 @@ public:
template<typename T>
void ASIMDLoadStoreSinglePost(uint32_t Op, uint32_t Q, uint32_t L, uint32_t R, uint32_t opcode, uint32_t S, uint32_t size,
ARMEmitter::Register rm, ARMEmitter::Register rn, T rt) {
LOGMAN_THROW_A_FMT(std::is_same_v<ARMEmitter::QRegister, T> || std::is_same_v<ARMEmitter::DRegister, T>, "Only supports 128-bit and "
"64-bit vector registers.");
LOGMAN_THROW_A_FMT((std::is_same_v<ARMEmitter::QRegister, T> || std::is_same_v<ARMEmitter::DRegister, T>), "Only supports 128-bit and "
"64-bit vector registers.");
uint32_t Instr = Op;
Instr |= Q << 30;
@@ -2136,7 +2132,370 @@ public:
}
// Memory copy/set
// TODO
void cpyfp(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0000, rs, rn, rd);
}
void cpyfm(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0000, rs, rn, rd);
}
void cpyfe(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0000, rs, rn, rd);
}
void cpyfpwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0001, rs, rn, rd);
}
void cpyfmwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0001, rs, rn, rd);
}
void cpyfewt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0001, rs, rn, rd);
}
void cpyfprt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0010, rs, rn, rd);
}
void cpyfmrt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0010, rs, rn, rd);
}
void cpyfert(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0010, rs, rn, rd);
}
void cpyfpt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0011, rs, rn, rd);
}
void cpyfmt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0011, rs, rn, rd);
}
void cpyfet(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0011, rs, rn, rd);
}
void cpyfpwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0100, rs, rn, rd);
}
void cpyfmwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0100, rs, rn, rd);
}
void cpyfewn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0100, rs, rn, rd);
}
void cpyfpwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0101, rs, rn, rd);
}
void cpyfmwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0101, rs, rn, rd);
}
void cpyfewtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0101, rs, rn, rd);
}
void cpyfprtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0110, rs, rn, rd);
}
void cpyfmrtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0110, rs, rn, rd);
}
void cpyfertwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0110, rs, rn, rd);
}
void cpyfptwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0111, rs, rn, rd);
}
void cpyfmtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0111, rs, rn, rd);
}
void cpyfetwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0111, rs, rn, rd);
}
void cpyfprn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1000, rs, rn, rd);
}
void cpyfmrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1000, rs, rn, rd);
}
void cpyfern(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1000, rs, rn, rd);
}
void cpyfpwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1001, rs, rn, rd);
}
void cpyfmwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1001, rs, rn, rd);
}
void cpyfewtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1001, rs, rn, rd);
}
void cpyfprtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1010, rs, rn, rd);
}
void cpyfmrtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1010, rs, rn, rd);
}
void cpyfertrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1010, rs, rn, rd);
}
void cpyfptrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1011, rs, rn, rd);
}
void cpyfmtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1011, rs, rn, rd);
}
void cpyfetrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1011, rs, rn, rd);
}
void cpyfpn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1100, rs, rn, rd);
}
void cpyfmn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1100, rs, rn, rd);
}
void cpyfen(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1100, rs, rn, rd);
}
void cpyfpwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1101, rs, rn, rd);
}
void cpyfmwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1101, rs, rn, rd);
}
void cpyfewtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1101, rs, rn, rd);
}
void cpyfprtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1110, rs, rn, rd);
}
void cpyfmrtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1110, rs, rn, rd);
}
void cpyfertn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1110, rs, rn, rd);
}
void cpyfptn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1111, rs, rn, rd);
}
void cpyfmtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1111, rs, rn, rd);
}
void cpyfetn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1111, rs, rn, rd);
}
void setp(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0000, rs, rn, rd);
}
void setm(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0100, rs, rn, rd);
}
void sete(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1000, rs, rn, rd);
}
void setpt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0001, rs, rn, rd);
}
void setmt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0101, rs, rn, rd);
}
void setet(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1001, rs, rn, rd);
}
void setpn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0010, rs, rn, rd);
}
void setmn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0110, rs, rn, rd);
}
void seten(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1010, rs, rn, rd);
}
void setptn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0011, rs, rn, rd);
}
void setmtn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0111, rs, rn, rd);
}
void setetn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1011, rs, rn, rd);
}
void cpyp(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0000, rs, rn, rd);
}
void cpym(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0000, rs, rn, rd);
}
void cpye(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0000, rs, rn, rd);
}
void cpypwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0001, rs, rn, rd);
}
void cpymwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0001, rs, rn, rd);
}
void cpyewt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0001, rs, rn, rd);
}
void cpyprt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0010, rs, rn, rd);
}
void cpymrt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0010, rs, rn, rd);
}
void cpyert(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0010, rs, rn, rd);
}
void cpypt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0011, rs, rn, rd);
}
void cpymt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0011, rs, rn, rd);
}
void cpyet(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0011, rs, rn, rd);
}
void cpypwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0100, rs, rn, rd);
}
void cpymwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0100, rs, rn, rd);
}
void cpyewn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0100, rs, rn, rd);
}
void cpypwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0101, rs, rn, rd);
}
void cpymwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0101, rs, rn, rd);
}
void cpyewtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0101, rs, rn, rd);
}
void cpyprtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0110, rs, rn, rd);
}
void cpymrtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0110, rs, rn, rd);
}
void cpyertwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0110, rs, rn, rd);
}
void cpyptwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0111, rs, rn, rd);
}
void cpymtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0111, rs, rn, rd);
}
void cpyetwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0111, rs, rn, rd);
}
void cpyprn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1000, rs, rn, rd);
}
void cpymrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1000, rs, rn, rd);
}
void cpyern(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1000, rs, rn, rd);
}
void cpypwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1001, rs, rn, rd);
}
void cpymwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1001, rs, rn, rd);
}
void cpyewtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1001, rs, rn, rd);
}
void cpyprtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1010, rs, rn, rd);
}
void cpymrtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1010, rs, rn, rd);
}
void cpyertrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1010, rs, rn, rd);
}
void cpyptrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1011, rs, rn, rd);
}
void cpymtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1011, rs, rn, rd);
}
void cpyetrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1011, rs, rn, rd);
}
void cpypn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1100, rs, rn, rd);
}
void cpymn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1100, rs, rn, rd);
}
void cpyen(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1100, rs, rn, rd);
}
void cpypwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1101, rs, rn, rd);
}
void cpymwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1101, rs, rn, rd);
}
void cpyewtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1101, rs, rn, rd);
}
void cpyprtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1110, rs, rn, rd);
}
void cpymrtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1110, rs, rn, rd);
}
void cpyertn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1110, rs, rn, rd);
}
void cpyptn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1111, rs, rn, rd);
}
void cpymtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1111, rs, rn, rd);
}
void cpyetn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1111, rs, rn, rd);
}
void setgp(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0000, rs, rn, rd);
}
void setgm(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0100, rs, rn, rd);
}
void setge(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1000, rs, rn, rd);
}
void setgpt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0001, rs, rn, rd);
}
void setgmt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0101, rs, rn, rd);
}
void setget(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1001, rs, rn, rd);
}
void setgpn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0010, rs, rn, rd);
}
void setgmn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0110, rs, rn, rd);
}
void setgen(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1010, rs, rn, rd);
}
void setgptn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0011, rs, rn, rd);
}
void setgmtn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0111, rs, rn, rd);
}
void setgetn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1011, rs, rn, rd);
}
// Loadstore no-allocate pair
void stnp(ARMEmitter::WRegister rt, ARMEmitter::WRegister rt2, ARMEmitter::Register rn, int32_t Imm) {
LOGMAN_THROW_A_FMT(Imm >= -256 && Imm <= 252 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -3819,7 +4178,12 @@ public:
}
// Loadstore PAC
// TODO
void ldraa(XRegister rt, XRegister rn, IndexType type, int32_t offset = 0) {
LoadStorePAC(0b11, 0, 0, offset, type, rn, rt);
}
void ldrab(XRegister rt, XRegister rn, IndexType type, int32_t offset = 0) {
LoadStorePAC(0b11, 0, 1, offset, type, rn, rt);
}
// Loadstore unsigned immediate
// Maximum values of unsigned immediate offsets for particular data sizes.
@@ -3968,6 +4332,20 @@ private:
dc32(Instr);
}
void MemoryCopyAndMemorySet(uint32_t sz, uint32_t o0, uint32_t op1, uint32_t op2, Register rs, Register rn, Register rd) {
uint32_t Instr = 0b0001'1001'0000'0000'0000'0100'0000'0000;
Instr |= sz << 30;
Instr |= o0 << 26;
Instr |= op1 << 22;
Instr |= rs.Idx() << 16;
Instr |= op2 << 12;
Instr |= rn.Idx() << 5;
Instr |= rd.Idx();
dc32(Instr);
}
// Loadstore no-allocate pair
template<typename T>
void LoadStoreNoAllocate(uint32_t Op, T rt, T rt2, ARMEmitter::Register rn, uint32_t Imm) {
@@ -4034,6 +4412,30 @@ private:
dc32(Instr);
}
void LoadStorePAC(uint32_t size, uint32_t VR, uint32_t M, int32_t imm, IndexType type, Register rn, Register rt) {
LOGMAN_THROW_A_FMT((imm % 8) == 0, "Immediate ({}) must be divisible by 8", imm);
LOGMAN_THROW_A_FMT(imm >= -4096 && imm <= 4088, "Immediate ({}) must be within [-4096, 4088]", imm);
LOGMAN_THROW_A_FMT(type == IndexType::OFFSET || type == IndexType::PRE, "PAC may only use offset or pre-indexed values");
// The immediate is scaled down in order to fit within the available 10 immediate bits.
const auto scaled_imm = static_cast<uint32_t>(imm / 8);
const auto imm9 = scaled_imm & 0b1'1111'1111;
const auto S = (scaled_imm >> 9) & 1;
const auto W = type == IndexType::OFFSET ? 0U : 1U;
uint32_t Instr = 0b0011'1000'0010'0000'0000'0100'0000'0000;
Instr |= size << 30;
Instr |= VR << 26;
Instr |= M << 23;
Instr |= S << 22;
Instr |= imm9 << 12;
Instr |= W << 11;
Instr |= rn.Idx() << 5;
Instr |= rt.Idx();
dc32(Instr);
}
// Loadstore unsigned immediate
template<typename T>
void LoadStoreUnsigned(uint32_t size, uint32_t V, uint32_t opc, T rt, Register rn, uint32_t Imm) {
+39 -39
View File
@@ -30,9 +30,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(Register) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<Register>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<Register>, "Needs to be standard");
static_assert(sizeof(Register) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<Register>);
static_assert(std::is_standard_layout_v<Register>);
/* 32-bit GPR register class.
* This class will imply a 32-bit register size being used.
@@ -58,9 +58,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(Register) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<Register>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<Register>, "Needs to be standard");
static_assert(sizeof(WRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<WRegister>);
static_assert(std::is_standard_layout_v<WRegister>);
/* 64-bit GPR register class.
* This class will imply a 64-bit register size being used.
@@ -86,9 +86,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(Register) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<Register>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<Register>, "Needs to be standard");
static_assert(sizeof(XRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<XRegister>);
static_assert(std::is_standard_layout_v<XRegister>);
inline constexpr WRegister Register::W() const {
return WRegister {Index};
@@ -283,9 +283,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(VRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<VRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<VRegister>, "Needs to be standard");
static_assert(sizeof(VRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<VRegister>);
static_assert(std::is_standard_layout_v<VRegister>);
/* 8-bit ASIMD register class
* This class implies 8-bit scalar register.
@@ -315,9 +315,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(BRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<BRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<BRegister>, "Needs to be standard");
static_assert(sizeof(BRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<BRegister>);
static_assert(std::is_standard_layout_v<BRegister>);
/* 16-bit ASIMD register class
* This class implies 16-bit scalar register.
@@ -347,9 +347,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(HRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<HRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<HRegister>, "Needs to be standard");
static_assert(sizeof(HRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<HRegister>);
static_assert(std::is_standard_layout_v<HRegister>);
/* 32-bit ASIMD register class
* This class implies 32-bit scalar register.
@@ -379,9 +379,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(SRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<SRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<SRegister>, "Needs to be standard");
static_assert(sizeof(SRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<SRegister>);
static_assert(std::is_standard_layout_v<SRegister>);
/* 64-bit ASIMD register class
* This class doesn't imply Vector or Scalar.
@@ -412,9 +412,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(DRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<DRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<DRegister>, "Needs to be standard");
static_assert(sizeof(DRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<DRegister>);
static_assert(std::is_standard_layout_v<DRegister>);
/* 128-bit ASIMD register class
* This class doesn't imply Vector or Scalar.
@@ -445,9 +445,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(QRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<QRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<QRegister>, "Needs to be standard");
static_assert(sizeof(QRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<QRegister>);
static_assert(std::is_standard_layout_v<QRegister>);
/* Unsized SVE register class.
* This class explicitly implies the instruction will operate using SVE.
@@ -474,9 +474,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(ZRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<ZRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<ZRegister>, "Needs to be standard");
static_assert(sizeof(ZRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<ZRegister>);
static_assert(std::is_standard_layout_v<ZRegister>);
// VRegister
inline constexpr BRegister VRegister::B() const {
@@ -919,9 +919,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(PRegister) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<PRegister>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<PRegister>, "Needs to be standard");
static_assert(sizeof(PRegister) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<PRegister>);
static_assert(std::is_standard_layout_v<PRegister>);
// Unsized predicate register for SVE with zeroing semantics.
class PRegisterZero {
@@ -947,9 +947,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(PRegisterZero) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<PRegisterZero>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<PRegisterZero>, "Needs to be standard");
static_assert(sizeof(PRegisterZero) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<PRegisterZero>);
static_assert(std::is_standard_layout_v<PRegisterZero>);
// Unsized predicate register for SVE with merging semantics.
class PRegisterMerge {
@@ -975,9 +975,9 @@ public:
private:
uint32_t Index;
};
static_assert(sizeof(PRegisterZero) == sizeof(uint32_t), "Needs to be uint32_t");
static_assert(std::is_trivial_v<PRegisterZero>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<PRegisterZero>, "Needs to be standard");
static_assert(sizeof(PRegisterMerge) == sizeof(uint32_t));
static_assert(std::is_trivially_copyable_v<PRegisterMerge>);
static_assert(std::is_standard_layout_v<PRegisterMerge>);
// PRegister
inline constexpr PRegisterZero PRegister::Zeroing() const {
+74 -38
View File
@@ -60,8 +60,7 @@ public:
}
void fcmla(SubRegSize size, ZRegister zda, PRegisterMerge pv, ZRegister zn, ZRegister zm, Rotation rot) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pv <= PReg::p7.Merging(), "fcmla can only use p0 to p7");
uint32_t Op = 0b0110'0100'0000'0000'0000'0000'0000'0000;
@@ -76,8 +75,7 @@ public:
}
void fcadd(SubRegSize size, ZRegister zd, PRegisterMerge pv, ZRegister zn, ZRegister zm, Rotation rot) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pv <= PReg::p7.Merging(), "fcadd can only use p0 to p7");
LOGMAN_THROW_A_FMT(rot == Rotation::ROTATE_90 || rot == Rotation::ROTATE_270, "fcadd rotation may only be 90 or 270 degrees");
LOGMAN_THROW_A_FMT(zd == zn, "fcadd zd and zn must be the same register");
@@ -815,16 +813,12 @@ public:
// SVE Integer Misc - Unpredicated
// SVE floating-point trig select coefficient
void ftssel(SubRegSize size, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "ftssel may only have "
"16-bit, 32-bit, or 64-bit "
"element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "ftssel may only use 16/32/64-bit element sizes");
SVEIntegerMiscUnpredicated(0b00, zm.Idx(), FEXCore::ToUnderlying(size), zd, zn);
}
// SVE floating-point exponential accelerator
void fexpa(SubRegSize size, ZRegister zd, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "fexpa may only have "
"16-bit, 32-bit, or 64-bit "
"element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "fexpa may only use 16/32/64-bit element sizes");
SVEIntegerMiscUnpredicated(0b10, 0b00000, FEXCore::ToUnderlying(size), zd, zn);
}
// SVE constructive prefix (unpredicated)
@@ -1503,9 +1497,9 @@ public:
}
// SVE broadcast floating-point immediate (unpredicated)
void fdup(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
LOGMAN_THROW_A_FMT(size == ARMEmitter::SubRegSize::i16Bit || size == ARMEmitter::SubRegSize::i32Bit || size == ARMEmitter::SubRegSize::i64Bit,
"Unsupported fmov size");
void fdup(SubRegSize size, ZRegister zd, float Value) {
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported fmov size");
uint32_t Imm {};
if (size == SubRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
@@ -1518,7 +1512,7 @@ public:
SVEBroadcastFloatImmUnpredicated(0b00, 0, Imm, size, zd);
}
void fmov(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
void fmov(SubRegSize size, ZRegister zd, float Value) {
fdup(size, zd, Value);
}
@@ -2134,11 +2128,31 @@ public:
// SVE2 Crypto Extensions
// SVE2 crypto unary operations
// XXX:
void aesimc(ZRegister zdn, ZRegister zn) {
SVE2CryptoUnaryOperation(1, zdn, zn);
}
void aesmc(ZRegister zdn, ZRegister zn) {
SVE2CryptoUnaryOperation(0, zdn, zn);
}
// SVE2 crypto destructive binary operations
// XXX:
void aese(ZRegister zdn, ZRegister zn, ZRegister zm) {
SVE2CryptoDestructiveBinaryOperation(0, 0, zdn, zn, zm);
}
void aesd(ZRegister zdn, ZRegister zn, ZRegister zm) {
SVE2CryptoDestructiveBinaryOperation(0, 1, zdn, zn, zm);
}
void sm4e(ZRegister zdn, ZRegister zn, ZRegister zm) {
SVE2CryptoDestructiveBinaryOperation(1, 0, zdn, zn, zm);
}
// SVE2 crypto constructive binary operations
// XXX:
void sm4ekey(ZRegister zd, ZRegister zn, ZRegister zm) {
SVE2CryptoConstructiveBinaryOperation(0, zd, zn, zm);
}
void rax1(ZRegister zd, ZRegister zn, ZRegister zm) {
SVE2CryptoConstructiveBinaryOperation(1, zd, zn, zm);
}
// SVE Floating Point Widening Multiply-Add - Indexed
// SVE BFloat16 floating-point dot product (indexed)
@@ -3381,8 +3395,8 @@ private:
}
void SVEBroadcastFloatImmPredicated(SubRegSize size, ZRegister zd, PRegister pg, float value) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported fcpy/fmov "
"size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported fcpy/fmov size");
uint32_t imm {};
if (size == SubRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
@@ -3558,7 +3572,7 @@ private:
// SVE2 floating-point pairwise operations
void SVEFloatPairwiseArithmetic(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(zd == zn, "zd needs to equal zn");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0100'0001'0000'1000'0000'0000'0000;
@@ -3572,7 +3586,7 @@ private:
// SVE floating-point arithmetic (unpredicated)
void SVEFloatArithmeticUnpredicated(uint32_t opc, SubRegSize size, ZRegister zm, ZRegister zn, ZRegister zd) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
uint32_t Instr = 0b0110'0101'0000'0000'0000'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -3680,7 +3694,7 @@ private:
// SVE floating-point arithmetic (predicated)
void SVEFloatArithmeticPredicated(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(zd == zn, "zn needs to equal zd");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0000'0000'1000'0000'0000'0000;
@@ -3708,9 +3722,7 @@ private:
}
void SVEFPRecursiveReduction(uint32_t opc, SubRegSize size, VRegister vd, PRegister pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "FP reduction operation can "
"only use 16-bit, 32-bit, "
"or 64-bit element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "FP reduction operation can only use 16/32/64-bit element sizes");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "FP reduction operation can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0000'0000'0010'0000'0000'0000;
@@ -3892,6 +3904,35 @@ private:
dc32(Instr);
}
void SVE2CryptoUnaryOperation(uint32_t op, ZRegister zdn, ZRegister zn) {
LOGMAN_THROW_A_FMT(zdn == zn, "zdn and zn must be the same register");
uint32_t Instr = 0b0100'0101'0010'0000'1110'0000'0000'0000;
Instr |= op << 10;
Instr |= zdn.Idx();
dc32(Instr);
}
void SVE2CryptoDestructiveBinaryOperation(uint32_t op, uint32_t o2, ZRegister zdn, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(zdn == zn, "zdn and zn must be the same register");
uint32_t Instr = 0b0100'0101'0010'0010'1110'0000'0000'0000;
Instr |= op << 16;
Instr |= o2 << 10;
Instr |= zm.Idx() << 5;
Instr |= zdn.Idx();
dc32(Instr);
}
void SVE2CryptoConstructiveBinaryOperation(uint32_t op, ZRegister zd, ZRegister zn, ZRegister zm) {
uint32_t Instr = 0b0100'0101'0010'0000'1111'0000'0000'0000;
Instr |= zm.Idx() << 16;
Instr |= op << 10;
Instr |= zn.Idx() << 5;
Instr |= zd.Idx();
dc32(Instr);
}
void SVE2BitwisePermute(SubRegSize size, uint32_t opc, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(size != SubRegSize::i128Bit, "Can't use 128-bit element size");
@@ -4063,7 +4104,7 @@ private:
// 0b111 - I - Current
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported size in {}", __func__);
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported size in {}", __func__);
uint32_t Instr = 0b0110'0101'0000'0000'1010'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4672,7 +4713,7 @@ private:
void SVEFloatUnary(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zn, ZRegister zd) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported size in {}", __func__);
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported size in {}", __func__);
uint32_t Instr = 0b0110'0101'0000'1100'1010'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4760,8 +4801,7 @@ private:
}
void SVEFPUnaryOpsUnpredicated(uint32_t opc, SubRegSize size, ZRegister zd, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
uint32_t Instr = 0b0110'0101'0000'1000'0011'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4772,8 +4812,7 @@ private:
}
void SVEFPSerialReductionPredicated(uint32_t opc, SubRegSize size, VRegister vd, PRegister pg, VRegister vn, ZRegister zm) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(vd == vn, "vn must be the same as vd");
@@ -4787,8 +4826,7 @@ private:
}
void SVEFPCompareWithZero(uint32_t eqlt, uint32_t ne, SubRegSize size, PRegister pd, PRegister pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0001'0000'0010'0000'0000'0000;
@@ -4803,8 +4841,7 @@ private:
void SVEFPMultiplyAdd(uint32_t opc, SubRegSize size, ZRegister zd, PRegister pg, ZRegister zn, ZRegister zm) {
// NOTE: opc also includes the op0 bit (bit 15) like op0:opc, since the fields are adjacent
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0010'0000'0000'0000'0000'0000;
@@ -4818,8 +4855,7 @@ private:
}
void SVEFPMultiplyAddIndexed(uint32_t op, SubRegSize size, ZRegister zda, ZRegister zn, ZRegister zm, uint32_t index) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT((size <= SubRegSize::i32Bit && zm <= ZReg::z7) || (size == SubRegSize::i64Bit && zm <= ZReg::z15),
"16-bit and 32-bit indexed variants may only use Zm between z0-z7\n"
"64-bit variants may only use Zm between z0-z15");
@@ -5051,7 +5087,7 @@ private:
const uint32_t element_size = SubRegSizeInBits(size);
if (is_left_shift) {
LOGMAN_THROW_A_FMT(shift >= 0 && shift < element_size, "Invalid left shift value ({}). Must be within [0, {}]", shift, element_size - 1);
LOGMAN_THROW_A_FMT(shift < element_size, "Invalid left shift value ({}). Must be within [0, {}]", shift, element_size - 1);
} else {
LOGMAN_THROW_A_FMT(shift > 0 && shift <= element_size, "Invalid right shift value ({}). Must be within [1, {}]", shift, element_size);
}
+111 -14
View File
@@ -27,8 +27,6 @@ struct EmitterOps : Emitter {
public:
// Advanced SIMD scalar copy
void dup(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Index) {
constexpr uint32_t Op = 0b0101'1110'0000'0000'0000'01 << 10;
const uint32_t SizeImm = FEXCore::ToUnderlying(size);
const uint32_t IndexShift = SizeImm + 1;
const uint32_t ElementSize = 1U << SizeImm;
@@ -38,10 +36,10 @@ public:
const uint32_t imm5 = (Index << IndexShift) | ElementSize;
ASIMDScalarCopy(Op, 1, imm5, 0b0000, rd, rn);
ASIMDScalarCopy(1, 1, imm5, 0b0000, rd, rn);
}
void mov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn, uint32_t Index) {
void mov(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Index) {
dup(size, rd, rn, Index);
}
@@ -137,7 +135,15 @@ public:
}
// Advanced SIMD scalar three same extra
// XXX:
void sqrdmlah(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i32Bit, "Only supports 16/32-bit");
ASIMDScalarThreeSameExtra(1, size, 0b0000, rm, rn, rd);
}
void sqrdmlsh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i32Bit, "Only supports 16/32-bit");
ASIMDScalarThreeSameExtra(1, size, 0b0001, rm, rn, rd);
}
// Advanced SIMD scalar two-register miscellaneous
void suqadd(ScalarRegSize size, VRegister rd, VRegister rn) {
ASIMDScalar2RegMisc(0, 0, size, 0b00011, rd, rn);
@@ -744,10 +750,51 @@ public:
const uint32_t immb = InvertedShift & 0b111;
ASIMDScalarShiftByImm(1, immh, immb, 0b10011, rd, rn);
}
// TODO: UCVTF, FCVTZU
// Advanced SIMD scalar x indexed element
// XXX:
//
void sqdmlal(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b0011, rm, rn, rd, index);
}
void sqdmlsl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b0111, rm, rn, rd, index);
}
void sqdmull(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b1011, rm, rn, rd, index);
}
void sqdmulh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b1100, rm, rn, rd, index);
}
void sqrdmulh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b1101, rm, rn, rd, index);
}
void fmla(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(0, size, 0b0001, rm, rn, rd, index);
}
void fmls(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(0, size, 0b0101, rm, rn, rd, index);
}
void fmul(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(0, size, 0b1001, rm, rn, rd, index);
}
void sqrdmlah(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(1, size, 0b1101, rm, rn, rd, index);
}
void sqrdmlsh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(1, size, 0b1111, rm, rn, rd, index);
}
void fmulx(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(1, size, 0b1001, rm, rn, rd, index);
}
// Floating-point data-processing (1 source)
void fmov(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b000000, rd, rn);
@@ -1233,10 +1280,10 @@ public:
private:
// Advanced SIMD scalar copy
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn) {
uint32_t Instr = Op;
void ASIMDScalarCopy(uint32_t Q, uint32_t b28, uint32_t imm5, uint32_t imm4, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0000'1110'0000'0000'0000'01U << 10;
Instr |= Q << 30;
Instr |= b28 << 28;
Instr |= imm5 << 16;
Instr |= imm4 << 11;
Instr |= Encode_rn(rn);
@@ -1269,7 +1316,17 @@ private:
}
// Advanced SIMD scalar three same extra
// XXX:
void ASIMDScalarThreeSameExtra(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rm, VRegister rn, VRegister rd) {
uint32_t Instr = 0b0101'1110'0000'0000'1000'0100'0000'0000;
Instr |= U << 29;
Instr |= FEXCore::ToUnderlying(size) << 22;
Instr |= rm.Idx() << 16;
Instr |= opcode << 11;
Instr |= rn.Idx() << 5;
Instr |= rd.Idx();
dc32(Instr);
}
// Advanced SIMD scalar two-register miscellaneous
void ASIMDScalar2RegMisc(uint32_t b20, uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0101'1110'0010'0000'0000'1000'0000'0000;
@@ -1283,8 +1340,6 @@ private:
dc32(Instr);
}
// Advanced SIMD scalar pairwise
// XXX:
// Advanced SIMD scalar three different
void ASIMD3RegDifferent(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0101'1110'0010'0000'0000'0000'0000'0000;
@@ -1321,8 +1376,50 @@ private:
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Advanced SIMD scalar x indexed element
// XXX:
void ASIMDScalarXIndexedElement(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rm, VRegister rn, VRegister rd, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i8Bit, "Scalar size must not be 8-bit");
[[maybe_unused]] const auto invalid_bound = 16U >> FEXCore::ToUnderlying(size);
LOGMAN_THROW_A_FMT(index < invalid_bound, "Index ({}) must be within [0-{}]", index, invalid_bound - 1);
uint32_t Instr = 0b0101'1111'0000'0000'0000'0000'0000'0000;
// FMUL/FMLA/FMLS indexed variants deal with size differently.
if (opcode == 0b0001 || opcode == 0b0101 || opcode == 0b1001) {
// Unlike other instructions in the group, 16-bit is encoded as zero
// and 32/64-bit are encoded with the top bit always set to one.
if (size != ScalarRegSize::i16Bit) {
Instr |= (0b10 | (FEXCore::ToUnderlying(size) & 1)) << 22;
}
} else {
Instr |= FEXCore::ToUnderlying(size) << 22;
}
uint32_t H = 0;
uint32_t LM = 0;
if (size == ScalarRegSize::i16Bit) {
LOGMAN_THROW_A_FMT(rm <= VReg::v15, "rm ({}) must be within [v0-v15]", rm.Idx());
H = (index >> 2) & 1;
LM = index & 0b11;
} else if (size == ScalarRegSize::i32Bit) {
H = (index >> 1) & 1;
LM = (index & 0b01) << 1;
} else {
H = index & 1;
}
Instr |= U << 29;
Instr |= LM << 20;
Instr |= rm.Idx() << 16;
Instr |= opcode << 12;
Instr |= H << 11;
Instr |= rn.Idx() << 5;
Instr |= rd.Idx();
dc32(Instr);
}
// Floating-point data-processing (1 source)
void Float1Source(uint32_t M, uint32_t S, uint32_t ptype, uint32_t opcode, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0001'1110'0010'0000'0100'0000'0000'0000;
+6
View File
@@ -13,6 +13,12 @@ struct EmitterOps : Emitter {
#endif
public:
// Reserved
void udf(uint32_t Imm) {
LOGMAN_THROW_A_FMT(Imm < 0x1'0000, "Immediate needs to be 16-bit");
dc32(Imm);
}
// System with result
// TODO: SYSL
// System Instruction
+5 -1
View File
@@ -224,9 +224,13 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// 11110s 2 UInt(s)
//
// So we 'or' (2 * -d) with our computed s to form imms.
if ((n != NULL) || (imm_s != NULL) || (imm_r != NULL)) {
if (n != nullptr) {
*n = out_n;
}
if (imm_s != nullptr) {
*imm_s = ((2 * -d) | (s - 1)) & 0x3f;
}
if (imm_r != nullptr) {
*imm_r = r;
}
File renamed without changes.
File renamed without changes.
File renamed without changes.
+3
View File
@@ -0,0 +1,3 @@
x86 and x86-64 Linux emulator
FEX allows you to run x86 applications on ARM64 Linux devices. It offers broad compatibility with both 32-bit and 64-bit binaries, and it can be used alongside Wine/Proton to play Windows games.
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
@@ -1,6 +1,6 @@
# This is a reference AArch64 cross compile script
# Pass in to cmake when building:
# eg: cmake -DCMAKE_TOOLCHAIN_FILE=../CMakeToolchains/AArch64.cmake ..
# eg: cmake --toolchain ../Data/CMake/toolchain_aarch64.cmake ..
if (NOT DEFINED ENV{SYSROOT})
message(FATAL_ERROR "Need to have SYSROOT environment variable set")
endif()
@@ -4,6 +4,7 @@ set(CMAKE_RC_COMPILER ${MINGW_TRIPLE}-windres)
set(CMAKE_C_COMPILER ${MINGW_TRIPLE}-clang)
set(CMAKE_CXX_COMPILER ${MINGW_TRIPLE}-clang++)
set(CMAKE_DLLTOOL ${MINGW_TRIPLE}-dlltool)
set(CMAKE_AR ${MINGW_TRIPLE}-ar)
# Compile everything as static to avoid requiring the MinGW runtime libraries, force page aligned sections so that
# debug symbols work correctly, and disable loop alignment to workaround an LLVM bug
File renamed without changes.
File renamed without changes.
File renamed without changes.
+31
View File
@@ -0,0 +1,31 @@
# --- Stage 1: Builder ---
FROM ubuntu:22.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
clang-13 llvm-13 nasm ninja-build pkg-config \
libcap-dev libglfw3-dev libepoxy-dev python3-dev libsdl2-dev \
python3 linux-headers-generic \
git qtbase5-dev qtdeclarative5-dev lld
RUN git clone --recurse-submodules https://github.com/FEX-Emu/FEX.git
WORKDIR /FEX
RUN mkdir build
ARG CC=clang-13
ARG CXX=clang++-13
RUN cmake -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_BUILD_TYPE=Release -DUSE_LINKER=lld -DENABLE_LTO=True -DBUILD_TESTS=False -DENABLE_ASSERTIONS=False -G Ninja .
RUN ninja
WORKDIR /FEX/build
# --- Stage 2: Runner ---
FROM builder as runner
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y \
libcap-dev libglfw3-dev libepoxy-dev
COPY --from=builder /FEX/Bin/* /usr/bin/
WORKDIR /
+35
View File
@@ -0,0 +1,35 @@
{ pkgs ? import <nixpkgs> { } }:
let
pkgsCross32 = pkgs.pkgsCross.gnu32;
pkgsCross64 = pkgs.pkgsCross.gnu64;
gcc32 = pkgs.writeText "toolchain_nix_gcc_x86_32.txt" ''
set(CMAKE_SYSTEM_PROCESSOR i686)
set(CMAKE_C_COMPILER ${pkgsCross32.buildPackages.gcc}/bin/i686-unknown-linux-gnu-gcc)
set(CMAKE_CXX_COMPILER ${pkgsCross32.buildPackages.gcc}/bin/i686-unknown-linux-gnu-g++)
'';
gcc64 = pkgs.writeText "toolchain_nix_gcc_x86_64.txt" ''
set(CMAKE_SYSTEM_PROCESSOR x86_64)
set(CMAKE_C_COMPILER ${pkgsCross64.buildPackages.gcc}/bin/x86_64-unknown-linux-gnu-gcc)
set(CMAKE_CXX_COMPILER ${pkgsCross64.buildPackages.gcc}/bin/x86_64-unknown-linux-gnu-g++)
'';
in
pkgs.mkShell {
buildInputs = [
pkgsCross64.buildPackages.clang
pkgsCross32.buildPackages.clang
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "toolchain32: ${gcc32}"
echo "toolchain64: ${gcc64}"
echo ""
echo "Use \$FEX_CMAKE_TOOLCHAINS to configure CMake."
fi
'';
FEX_CMAKE_TOOLCHAINS = "-DX86_32_TOOLCHAIN_FILE=${gcc32} -DX86_64_TOOLCHAIN_FILE=${gcc64}";
}
+83
View File
@@ -0,0 +1,83 @@
{ pkgs ? import <nixpkgs> { } }:
let
pkgsCross32 = pkgs.pkgsCross.gnu32;
pkgsCross64 = pkgs.pkgsCross.gnu64;
devRootFS = pkgs.buildEnv {
name = "fex-dev-rootfs";
paths = [
pkgsCross64.stdenv.cc.libc_dev
pkgsCross32.stdenv.cc.libc_dev
pkgsCross64.stdenv.cc.cc
pkgsCross32.stdenv.cc.cc
pkgs.alsa-lib.dev
pkgs.libdrm.dev
pkgs.libGL.dev
pkgs.wayland.dev
pkgs.xorg.libX11.dev
pkgs.xorg.libxcb.dev
pkgs.xorg.libXrandr.dev
pkgs.xorg.libXrender.dev
pkgs.xorg.xorgproto
];
ignoreCollisions = true;
pathsToLink = [
"/include"
"/lib"
];
postBuild = ''
mkdir -p $out/usr
ln -s $out/include $out/usr/
'';
};
toolchain32 = pkgs.writeText "toolchain_nix_x86_32.txt" ''
set(CMAKE_EXE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_MODULE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SYSTEM_PROCESSOR i686)
set(CMAKE_C_COMPILER clang)
set(CMAKE_CXX_COMPILER clang++)
set(CMAKE_C_COMPILER ${pkgsCross32.buildPackages.clang}/bin/i686-unknown-linux-gnu-clang)
set(CMAKE_CXX_COMPILER ${pkgsCross32.buildPackages.clang}/bin/i686-unknown-linux-gnu-clang++)
set(CLANG_FLAGS "-nodefaultlibs -nostartfiles -lstdc++ -target i686-linux-gnu -msse2 -mfpmath=sse --sysroot=${devRootFS} -iwithsysroot/include")
set(CMAKE_C_FLAGS "''${CMAKE_C_FLAGS} ''${CLANG_FLAGS}")
set(CMAKE_CXX_FLAGS "''${CMAKE_CXX_FLAGS} ''${CLANG_FLAGS}")
'';
toolchain64 = pkgs.writeText "toolchain_nix_x86_64.txt" ''
set(CMAKE_EXE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_MODULE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SYSTEM_PROCESSOR x86_64)
set(CMAKE_C_COMPILER clang)
set(CMAKE_CXX_COMPILER clang++)
set(CMAKE_C_COMPILER ${pkgsCross64.buildPackages.clang}/bin/x86_64-unknown-linux-gnu-clang)
set(CMAKE_CXX_COMPILER ${pkgsCross64.buildPackages.clang}/bin/x86_64-unknown-linux-gnu-clang++)
set(CLANG_FLAGS "-nodefaultlibs -nostartfiles -lstdc++ -target x86_64-linux-gnu --sysroot=${devRootFS} -iwithsysroot/usr/include")
set(CMAKE_C_FLAGS "''${CMAKE_C_FLAGS} ''${CLANG_FLAGS}")
set(CMAKE_CXX_FLAGS "''${CMAKE_CXX_FLAGS} ''${CLANG_FLAGS}")
'';
in
pkgs.mkShell {
buildInputs = [
pkgsCross64.buildPackages.clang
pkgsCross32.buildPackages.clang
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "Set up dev RootFS at ${devRootFS}"
echo "toolchain32: ${toolchain32}"
echo "toolchain64: ${toolchain64}"
echo ""
echo "Use \$FEX_CMAKE_TOOLCHAINS to configure CMake."
fi
'';
FEX_CMAKE_TOOLCHAINS = "-DX86_32_TOOLCHAIN_FILE=${toolchain32} -DX86_64_TOOLCHAIN_FILE=${toolchain64} -DX86_DEV_ROOTFS=${devRootFS}";
ROOTFS = "${devRootFS}";
}
+52
View File
@@ -0,0 +1,52 @@
{ pkgs ? import <nixpkgs> { } }:
let
toolchain = pkgs.fetchzip {
url = "https://github.com/bylaws/llvm-mingw/releases/download/20250305/llvm-mingw-20250305-ucrt-ubuntu-20.04-aarch64.tar.xz";
sha256 = "sha256-cA03/ab9O61eO9+S2JzIXD4V0HzTXK5/AYyxW2d73Po=";
};
cmakeToolchainFile = pkgs.substitute {
# Use absolute paths that are discoverable outside of the nix shell
src = ../../CMake/toolchain_mingw.cmake;
substitutions = ["--replace-fail" "\${MINGW_TRIPLE}-" "${toolchain}/bin/\${MINGW_TRIPLE}-"];
};
mesonCrossFile = pkgs.writeText "crossfile_llvm_mingw.txt" ''
[binaries]
ar = '${toolchain}/bin/arm64ec-w64-mingw32-ar'
c = '${toolchain}/bin/arm64ec-w64-mingw32-gcc'
cpp = '${toolchain}/bin/arm64ec-w64-mingw32-g++'
ld = '${toolchain}/bin/arm64ec-w64-mingw32-ld'
windres = '${toolchain}/bin/arm64ec-w64-mingw32-windres'
strip = '${toolchain}/bin/strip'
widl = '${toolchain}/bin/arm64ec-w64-mingw32-widl'
pkgconfig = 'aarch64-linux-gnu-pkg-config'
[host_machine]
system = 'windows'
cpu_family = 'aarch64'
cpu = 'aarch64'
endian = 'little'
'';
in
pkgs.mkShell {
buildInputs = [
toolchain
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "llvm-mingw set up at ${toolchain}."
echo ""
echo "To configure DXVK/vkd3d-proton: meson setup \$FEX_MESON_CROSSFILE"
echo ""
echo "To configure 32-bit FEX build: cmake \$FEX_CMAKE_TOOLCHAIN_WOW64"
echo "To configure 64-bit FEX build: cmake \$FEX_CMAKE_TOOLCHAIN_ARM64EC"
fi
'';
# E.g. cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False
FEX_CMAKE_TOOLCHAIN_ARM64EC = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_CMAKE_TOOLCHAIN_WOW64 = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_MESON_CROSSFILE = "--cross-file ${mesonCrossFile}";
}
+21
View File
@@ -0,0 +1,21 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash WineOnArm/shell.nix
# Helper script to configure CMake for building FEX as library for emulation
# of 32-bit applications in Wine/Proton.
# The required cross-toolchains will be set up and managed by nix.
if [ $# -eq 0 ]
then
echo "Expected CMake argument list"
exit 1
fi
if [ -f CMakeCache.txt ]
then
echo "Expected empty build folder"
exit 1
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_WOW64 -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
+21
View File
@@ -0,0 +1,21 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash WineOnArm/shell.nix
# Helper script to configure CMake for building FEX as library for emulation
# of 64-bit applications in Wine/Proton
# Nix is used to install and manage the required cross-toolchains.
if [ $# -eq 0 ]
then
echo "Expected CMake argument list"
exit 1
fi
if [ -f CMakeCache.txt ]
then
echo "Expected empty build folder"
exit 1
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
+17
View File
@@ -0,0 +1,17 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash FEXLinuxTests/shell.nix
# Helper script to configure CMake for building FEXLinuxTests.
# Nix is used to install and manage the required cross-toolchains.
if [ ! -f CMakeCache.txt ]
then
echo "Must be run from a pre-configured CMake build folder"
exit 1
fi
# Remove previous build to ensure the new toolchain is applied
rm -rf unittests/FEXLinuxTests
set -o xtrace
cmake . $FEX_CMAKE_TOOLCHAINS -DBUILD_TESTS=ON -DBUILD_FEX_LINUX_TESTS=ON
+22
View File
@@ -0,0 +1,22 @@
# Helper script to configure CMake for library forwarding in FEX.
# Nix is used to install and manage the required cross-toolchains.
if [ ! -f CMakeCache.txt ]
then
echo "Must be run from a pre-configured CMake build folder"
exit 1
fi
# Remove previous build to ensure the new toolchain is applied
rm -rf guest-libs guest-libs-32 Guest Guest_32
# Set clang executable path manually since the one from the nix store
# will be picked up otherwise
CLANG_EXEC_PATH=""
if ! grep -q CLANG_EXEC_PATH CMakeCache.txt
then
CLANG_EXEC_PATH="-DCLANG_EXEC_PATH=`which clang`"
fi
nix-shell `dirname -- "$0"`/LibraryForwarding/shell.nix \
--run "set -o xtrace; cmake . \$FEX_CMAKE_TOOLCHAINS -DBUILD_THUNKS=ON $CLANG_EXEC_PATH; set +o xtrace"
-31
View File
@@ -1,31 +0,0 @@
# --- Stage 1: Builder ---
FROM ubuntu:20.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
clang-10 llvm-10 nasm ninja-build pkg-config \
libcap-dev libglfw3-dev libepoxy-dev python3-dev libsdl2-dev \
python3 linux-headers-generic \
git
RUN git clone --recurse-submodules https://github.com/FEX-Emu/FEX.git
CMD [ "mkdir /opt/FEX/build" ]
WORKDIR /opt/FEX/build
ARG CC=clang-10
ARG CXX=clang++-10
RUN cmake -G Ninja .. -DCMAKE_BUILD_TYPE=Release
RUN ninja
# --- Stage 2: Runner ---
FROM ubuntu:20.04
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y \
libcap-dev libglfw3-dev libepoxy-dev
COPY --from=builder /opt/FEX/build/Bin/* /usr/bin/
WORKDIR /root
+1
View File
@@ -0,0 +1 @@
DisableFormat: true
+98
View File
@@ -0,0 +1,98 @@
set (SRCS
# F80 support
src/extF80_add.c
src/extF80_div.c
src/extF80_sub.c
src/extF80_mul.c
src/extF80_rem.c
src/extF80_sqrt.c
src/extF80_le.c
src/extF80_to_i32.c
src/extF80_to_i64.c
src/extF80_to_ui64.c
src/extF80_to_f32.c
src/extF80_to_f64.c
src/i32_to_extF80.c
src/ui64_to_extF80.c
src/extF80_to_f128.c
src/f128_to_extF80.c
# F128 support
src/f128_add.c
src/f128_div.c
src/f128_eq.c
src/f128_eq_signaling.c
src/f128_isSignalingNaN.c
src/f128_le.c
src/f128_le_quiet.c
src/f128_lt.c
src/f128_lt_quiet.c
src/f128_mulAdd.c
src/f128_mul.c
src/f128_rem.c
src/f128_sqrt.c
src/f128_sub.c
src/f128_to_f16.c
src/f128_to_f32.c
src/f128_to_f64.c
src/f128_to_i32.c
src/f128_to_i64.c
src/f128_to_ui32.c
src/f128_to_ui64.c
src/s_addMagsF128.c
src/s_subMagsF128.c
src/s_normRoundPackToF128.c
src/s_roundPackToF128.c
src/s_propagateNaNF128UI.c
# Conversion
src/f32_to_f128.c
src/i32_to_f128.c
src/s_roundToUI64.c
src/s_f128UIToCommonNaN.c
src/s_commonNaNToF128UI.c
src/s_normSubnormalF128Sig.c
src/s_roundToI32.c
src/s_roundToI64.c
src/s_roundPackToF32.c
src/s_addMagsExtF80.c
src/s_extF80UIToCommonNaN.c
src/s_commonNaNToF32UI.c
src/s_commonNaNToF64UI.c
src/s_roundPackToF64.c
src/s_propagateNaNExtF80UI.c
src/s_roundPackToExtF80.c
src/s_normSubnormalExtF80Sig.c
src/s_subMagsExtF80.c
src/s_shiftRightJam128.c
src/s_shiftRightJam128Extra.c
src/s_normRoundPackToExtF80.c
src/s_approxRecip_1Ks.c
src/s_approxRecipSqrt32_1.c
src/s_approxRecipSqrt_1Ks.c
src/softfloat_raiseFlags.c
src/f64_to_extF80.c
src/s_commonNaNToExtF80UI.c
src/s_normSubnormalF64Sig.c
src/s_f64UIToCommonNaN.c
src/extF80_roundToInt.c
src/extF80_eq.c
src/extF80_lt.c
src/f32_to_extF80.c
src/s_normSubnormalF32Sig.c
src/s_f32UIToCommonNaN.c)
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=__attribute__((preserve_all));-DFEXCORE_HAS_PRESERVE_ALL_ATTR=1")
else()
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=;-DFEXCORE_HAS_PRESERVE_ALL_ATTR=0")
endif()
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ=1;-DINLINE=static inline;-DINLINE_LEVEL=4;-DSOFTFLOAT_FAST_INT64=1;-DSOFTFLOAT_FAST_DIV32TO16=1;-DSOFTFLOAT_FAST_DIV64TO32=1")
add_library(softfloat_3e STATIC ${SRCS})
target_include_directories(softfloat_3e PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}/include/)
target_include_directories(softfloat_3e PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}/include/SoftFloat-3e/)
target_compile_definitions(softfloat_3e PUBLIC ${DEFINES})
@@ -149,7 +149,7 @@ float64_t f32_to_f64( float32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( struct softfloat_state *, float32_t );
float128_t f32_to_f128( float32_t );
float128_t f32_to_f128( struct softfloat_state *, float32_t );
#endif
void f32_to_extF80M( float32_t, extFloat80_t * );
void f32_to_f128M( float32_t, float128_t * );
@@ -243,13 +243,17 @@ FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_le( extFloat80_t, extFloat80_t );
bool extF80_le( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_eq_signaling( extFloat80_t, extFloat80_t );
bool extF80_le_quiet( extFloat80_t, extFloat80_t );
bool extF80_lt_quiet( extFloat80_t, extFloat80_t );
bool extF80_isSignalingNaN( extFloat80_t );
static inline extFloat80_t extF80_complement_sign(extFloat80_t a) {
a.signExp ^= 1ULL << 15;
return a;
}
#endif
uint_fast32_t extF80M_to_ui32( const extFloat80_t *, uint_fast8_t, bool );
uint_fast64_t extF80M_to_ui64( const extFloat80_t *, uint_fast8_t, bool );
@@ -284,34 +288,38 @@ bool extF80M_isSignalingNaN( const extFloat80_t * );
| 128-bit (quadruple-precision) floating-point operations.
*----------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_FAST_INT64
uint_fast32_t f128_to_ui32( float128_t, uint_fast8_t, bool );
uint_fast64_t f128_to_ui64( float128_t, uint_fast8_t, bool );
int_fast32_t f128_to_i32( float128_t, uint_fast8_t, bool );
int_fast64_t f128_to_i64( float128_t, uint_fast8_t, bool );
uint_fast32_t f128_to_ui32( struct softfloat_state *, float128_t, uint_fast8_t, bool );
uint_fast64_t f128_to_ui64( struct softfloat_state *, float128_t, uint_fast8_t, bool );
int_fast32_t f128_to_i32( struct softfloat_state *, float128_t, uint_fast8_t, bool );
int_fast64_t f128_to_i64( struct softfloat_state *, float128_t, uint_fast8_t, bool );
uint_fast32_t f128_to_ui32_r_minMag( float128_t, bool );
uint_fast64_t f128_to_ui64_r_minMag( float128_t, bool );
int_fast32_t f128_to_i32_r_minMag( float128_t, bool );
int_fast64_t f128_to_i64_r_minMag( float128_t, bool );
float16_t f128_to_f16( float128_t );
float32_t f128_to_f32( float128_t );
float64_t f128_to_f64( float128_t );
float16_t f128_to_f16( struct softfloat_state *, float128_t );
float32_t f128_to_f32( struct softfloat_state *, float128_t );
float64_t f128_to_f64( struct softfloat_state *, float128_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( struct softfloat_state *, float128_t );
float128_t f128_roundToInt( float128_t, uint_fast8_t, bool );
float128_t f128_add( float128_t, float128_t );
float128_t f128_sub( float128_t, float128_t );
float128_t f128_mul( float128_t, float128_t );
float128_t f128_mulAdd( float128_t, float128_t, float128_t );
float128_t f128_div( float128_t, float128_t );
float128_t f128_rem( float128_t, float128_t );
float128_t f128_sqrt( float128_t );
bool f128_eq( float128_t, float128_t );
bool f128_le( float128_t, float128_t );
bool f128_lt( float128_t, float128_t );
bool f128_eq_signaling( float128_t, float128_t );
bool f128_le_quiet( float128_t, float128_t );
bool f128_lt_quiet( float128_t, float128_t );
float128_t f128_add( struct softfloat_state *, float128_t, float128_t );
float128_t f128_sub( struct softfloat_state *, float128_t, float128_t );
float128_t f128_mul( struct softfloat_state *, float128_t, float128_t );
float128_t f128_mulAdd( struct softfloat_state *, float128_t, float128_t, float128_t );
float128_t f128_div( struct softfloat_state *, float128_t, float128_t );
float128_t f128_rem( struct softfloat_state *, float128_t, float128_t );
float128_t f128_sqrt( struct softfloat_state *, float128_t );
bool f128_eq( struct softfloat_state *, float128_t, float128_t );
bool f128_le( struct softfloat_state *, float128_t, float128_t );
bool f128_lt( struct softfloat_state *, float128_t, float128_t );
bool f128_eq_signaling( struct softfloat_state *, float128_t, float128_t );
bool f128_le_quiet( struct softfloat_state *, float128_t, float128_t );
bool f128_lt_quiet( struct softfloat_state *, float128_t, float128_t );
bool f128_isSignalingNaN( float128_t );
static inline float128_t f128_complement_sign(float128_t a) {
a.v[1] ^= 1ULL << 63;
return a;
}
#endif
uint_fast32_t f128M_to_ui32( const float128_t *, uint_fast8_t, bool );
uint_fast64_t f128M_to_ui64( const float128_t *, uint_fast8_t, bool );
File renamed without changes.
File renamed without changes.
File renamed without changes.
+73
View File
@@ -0,0 +1,73 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool extF80_le( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
uint_fast64_t uiA0;
union { struct extFloat80M s; extFloat80_t f; } uB;
uint_fast16_t uiB64;
uint_fast64_t uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.s.signExp;
uiA0 = uA.s.signif;
uB.f = b;
uiB64 = uB.s.signExp;
uiB0 = uB.s.signif;
if ( isNaNExtF80UI( uiA64, uiA0 ) || isNaNExtF80UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signExtF80UI64( uiA64 );
signB = signExtF80UI64( uiB64 );
return
(signA != signB)
? signA || ! (((uiA64 | uiB64) & 0x7FFF) | uiA0 | uiB0)
: ((uiA64 == uiB64) && (uiA0 == uiB0))
|| (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t f128_add( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
#if ! defined INLINE_LEVEL || (INLINE_LEVEL < 2)
float128_t
(*magsFuncPtr)(
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
#endif
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
#if defined INLINE_LEVEL && (2 <= INLINE_LEVEL)
if ( signA == signB ) {
return softfloat_addMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
} else {
return softfloat_subMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
}
#else
magsFuncPtr =
(signA == signB) ? softfloat_addMagsF128 : softfloat_subMagsF128;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
#endif
}
+199
View File
@@ -0,0 +1,199 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_div( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
int_fast32_t expB;
struct uint128 sigB;
bool signZ;
struct exp32_sig128 normExpSig;
int_fast32_t expZ;
struct uint128 rem;
uint_fast32_t recip32;
int ix;
uint_fast64_t q64;
uint_fast32_t q;
struct uint128 term;
uint_fast32_t qs[3];
uint_fast64_t sigZExtra;
struct uint128 sigZ, uiZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
signZ = signA ^ signB;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 ) goto propagateNaN;
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
goto invalid;
}
goto infinity;
}
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
goto zero;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expB ) {
if ( ! (sigB.v64 | sigB.v0) ) {
if ( ! (expA | sigA.v64 | sigA.v0) ) goto invalid;
softfloat_raiseFlags( state, softfloat_flag_infinite );
goto infinity;
}
normExpSig = softfloat_normSubnormalF128Sig( sigB.v64, sigB.v0 );
expB = normExpSig.exp;
sigB = normExpSig.sig;
}
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) goto zero;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
expZ = expA - expB + 0x3FFE;
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sigB.v64 |= UINT64_C( 0x0001000000000000 );
rem = sigA;
if ( softfloat_lt128( sigA.v64, sigA.v0, sigB.v64, sigB.v0 ) ) {
--expZ;
rem = softfloat_add128( sigA.v64, sigA.v0, sigA.v64, sigA.v0 );
}
recip32 = softfloat_approxRecip32_1( sigB.v64>>17 );
ix = 3;
for (;;) {
q64 = (uint_fast64_t) (uint32_t) (rem.v64>>19) * recip32;
q = (q64 + 0x80000000)>>32;
--ix;
if ( ix < 0 ) break;
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
--q;
rem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
qs[ix] = q;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ((q + 1) & 7) < 2 ) {
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
--q;
rem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
} else if ( softfloat_le128( sigB.v64, sigB.v0, rem.v64, rem.v0 ) ) {
++q;
rem = softfloat_sub128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
if ( rem.v64 | rem.v0 ) q |= 1;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
sigZExtra = (uint64_t) ((uint_fast64_t) q<<60);
term = softfloat_shortShiftLeft128( 0, qs[1], 54 );
sigZ =
softfloat_add128(
(uint_fast64_t) qs[2]<<19, ((uint_fast64_t) qs[0]<<25) + (q>>4),
term.v64, term.v0
);
return
softfloat_roundPackToF128( state, signZ, expZ, sigZ.v64, sigZ.v0, sigZExtra );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
infinity:
uiZ.v64 = packToF128UI64( signZ, 0x7FFF, 0 );
goto uiZ0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
zero:
uiZ.v64 = packToF128UI64( signZ, 0, 0 );
uiZ0:
uiZ.v0 = 0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+73
View File
@@ -0,0 +1,73 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_eq( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
if (
softfloat_isSigNaNF128UI( uiA64, uiA0 )
|| softfloat_isSigNaNF128UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
return
(uiA0 == uiB0)
&& ( (uiA64 == uiB64)
|| (! uiA0 && ! ((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF )))
);
}
+67
View File
@@ -0,0 +1,67 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
bool f128_eq_signaling( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
return
(uiA0 == uiB0)
&& ( (uiA64 == uiB64)
|| (! uiA0 && ! ((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF )))
);
}
+51
View File
@@ -0,0 +1,51 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_isSignalingNaN( float128_t a )
{
union ui128_f128 uA;
uA.f = a;
return softfloat_isSigNaNF128UI( uA.ui.v64, uA.ui.v0 );
}
+72
View File
@@ -0,0 +1,72 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
bool f128_le( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
|| ! (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 == uiB64) && (uiA0 == uiB0))
|| (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_le_quiet( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
if (
softfloat_isSigNaNF128UI( uiA64, uiA0 )
|| softfloat_isSigNaNF128UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
|| ! (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 == uiB64) && (uiA0 == uiB0))
|| (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+72
View File
@@ -0,0 +1,72 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
bool f128_lt( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
&& (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 != uiB64) || (uiA0 != uiB0))
&& (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_lt_quiet( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
if (
softfloat_isSigNaNF128UI( uiA64, uiA0 )
|| softfloat_isSigNaNF128UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
&& (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 != uiB64) || (uiA0 != uiB0))
&& (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+163
View File
@@ -0,0 +1,163 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_mul( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
int_fast32_t expB;
struct uint128 sigB;
bool signZ;
uint_fast64_t magBits;
struct exp32_sig128 normExpSig;
int_fast32_t expZ;
uint64_t sig256Z[4];
uint_fast64_t sigZExtra;
struct uint128 sigZ;
struct uint128_extra sig128Extra;
struct uint128 uiZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
signZ = signA ^ signB;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if (
(sigA.v64 | sigA.v0) || ((expB == 0x7FFF) && (sigB.v64 | sigB.v0))
) {
goto propagateNaN;
}
magBits = expB | sigB.v64 | sigB.v0;
goto infArg;
}
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
magBits = expA | sigA.v64 | sigA.v0;
goto infArg;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) goto zero;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
if ( ! expB ) {
if ( ! (sigB.v64 | sigB.v0) ) goto zero;
normExpSig = softfloat_normSubnormalF128Sig( sigB.v64, sigB.v0 );
expB = normExpSig.exp;
sigB = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
expZ = expA + expB - 0x4000;
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sigB = softfloat_shortShiftLeft128( sigB.v64, sigB.v0, 16 );
softfloat_mul128To256M( sigA.v64, sigA.v0, sigB.v64, sigB.v0, sig256Z );
sigZExtra = sig256Z[indexWord( 4, 1 )] | (sig256Z[indexWord( 4, 0 )] != 0);
sigZ =
softfloat_add128(
sig256Z[indexWord( 4, 3 )], sig256Z[indexWord( 4, 2 )],
sigA.v64, sigA.v0
);
if ( UINT64_C( 0x0002000000000000 ) <= sigZ.v64 ) {
++expZ;
sig128Extra =
softfloat_shortShiftRightJam128Extra(
sigZ.v64, sigZ.v0, sigZExtra, 1 );
sigZ = sig128Extra.v;
sigZExtra = sig128Extra.extra;
}
return
softfloat_roundPackToF128( state, signZ, expZ, sigZ.v64, sigZ.v0, sigZExtra );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
infArg:
if ( ! magBits ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
goto uiZ;
}
uiZ.v64 = packToF128UI64( signZ, 0x7FFF, 0 );
goto uiZ0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
zero:
uiZ.v64 = packToF128UI64( signZ, 0, 0 );
uiZ0:
uiZ.v0 = 0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+63
View File
@@ -0,0 +1,63 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t f128_mulAdd( struct softfloat_state *state, float128_t a, float128_t b, float128_t c )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
union ui128_f128 uC;
uint_fast64_t uiC64, uiC0;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
uC.f = c;
uiC64 = uC.ui.v64;
uiC0 = uC.ui.v0;
return softfloat_mulAddF128( uiA64, uiA0, uiB64, uiB0, uiC64, uiC0, 0 );
}
+190
View File
@@ -0,0 +1,190 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_rem( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
int_fast32_t expB;
struct uint128 sigB;
struct exp32_sig128 normExpSig;
struct uint128 rem;
int_fast32_t expDiff;
uint_fast32_t q, recip32;
uint_fast64_t q64;
struct uint128 term, altRem, meanRem;
bool signRem;
struct uint128 uiZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if (
(sigA.v64 | sigA.v0) || ((expB == 0x7FFF) && (sigB.v64 | sigB.v0))
) {
goto propagateNaN;
}
goto invalid;
}
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
return a;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expB ) {
if ( ! (sigB.v64 | sigB.v0) ) goto invalid;
normExpSig = softfloat_normSubnormalF128Sig( sigB.v64, sigB.v0 );
expB = normExpSig.exp;
sigB = normExpSig.sig;
}
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) return a;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sigB.v64 |= UINT64_C( 0x0001000000000000 );
rem = sigA;
expDiff = expA - expB;
if ( expDiff < 1 ) {
if ( expDiff < -1 ) return a;
if ( expDiff ) {
--expB;
sigB = softfloat_add128( sigB.v64, sigB.v0, sigB.v64, sigB.v0 );
q = 0;
} else {
q = softfloat_le128( sigB.v64, sigB.v0, rem.v64, rem.v0 );
if ( q ) {
rem = softfloat_sub128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
}
} else {
recip32 = softfloat_approxRecip32_1( sigB.v64>>17 );
expDiff -= 30;
for (;;) {
q64 = (uint_fast64_t) (uint32_t) (rem.v64>>19) * recip32;
if ( expDiff < 0 ) break;
q = (q64 + 0x80000000)>>32;
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
rem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
expDiff -= 29;
}
/*--------------------------------------------------------------------
| (`expDiff' cannot be less than -29 here.)
*--------------------------------------------------------------------*/
q = (uint32_t) (q64>>32)>>(~expDiff & 31);
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, expDiff + 30 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
altRem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
goto selectRem;
}
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
do {
altRem = rem;
++q;
rem = softfloat_sub128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
} while ( ! (rem.v64 & UINT64_C( 0x8000000000000000 )) );
selectRem:
meanRem = softfloat_add128( rem.v64, rem.v0, altRem.v64, altRem.v0 );
if (
(meanRem.v64 & UINT64_C( 0x8000000000000000 ))
|| (! (meanRem.v64 | meanRem.v0) && (q & 1))
) {
rem = altRem;
}
signRem = signA;
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
signRem = ! signRem;
rem = softfloat_sub128( 0, 0, rem.v64, rem.v0 );
}
return softfloat_normRoundPackToF128( state, signRem, expB - 1, rem.v64, rem.v0 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+201
View File
@@ -0,0 +1,201 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_sqrt( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA, uiZ;
struct exp32_sig128 normExpSig;
int_fast32_t expZ;
uint_fast32_t sig32A, recipSqrt32, sig32Z;
struct uint128 rem;
uint32_t qs[3];
uint_fast32_t q;
uint_fast64_t x64, sig64Z;
struct uint128 y, term;
uint_fast64_t sigZExtra;
struct uint128 sigZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 ) {
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, 0, 0 );
goto uiZ;
}
if ( ! signA ) return a;
goto invalid;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( signA ) {
if ( ! (expA | sigA.v64 | sigA.v0) ) return a;
goto invalid;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) return a;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
/*------------------------------------------------------------------------
| (`sig32Z' is guaranteed to be a lower bound on the square root of
| `sig32A', which makes `sig32Z' also a lower bound on the square root of
| `sigA'.)
*------------------------------------------------------------------------*/
expZ = ((expA - 0x3FFF)>>1) + 0x3FFE;
expA &= 1;
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sig32A = sigA.v64>>17;
recipSqrt32 = softfloat_approxRecipSqrt32_1( expA, sig32A );
sig32Z = ((uint_fast64_t) sig32A * recipSqrt32)>>32;
if ( expA ) {
sig32Z >>= 1;
rem = softfloat_shortShiftLeft128( sigA.v64, sigA.v0, 12 );
} else {
rem = softfloat_shortShiftLeft128( sigA.v64, sigA.v0, 13 );
}
qs[2] = sig32Z;
rem.v64 -= (uint_fast64_t) sig32Z * sig32Z;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
q = ((uint32_t) (rem.v64>>2) * (uint_fast64_t) recipSqrt32)>>32;
x64 = (uint_fast64_t) sig32Z<<32;
sig64Z = x64 + ((uint_fast64_t) q<<3);
y = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
/*------------------------------------------------------------------------
| (Repeating this loop is a rare occurrence.)
*------------------------------------------------------------------------*/
for (;;) {
term = softfloat_mul64ByShifted32To128( x64 + sig64Z, q );
rem = softfloat_sub128( y.v64, y.v0, term.v64, term.v0 );
if ( ! (rem.v64 & UINT64_C( 0x8000000000000000 )) ) break;
--q;
sig64Z -= 1<<3;
}
qs[1] = q;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
q = ((rem.v64>>2) * recipSqrt32)>>32;
y = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
sig64Z <<= 1;
/*------------------------------------------------------------------------
| (Repeating this loop is a rare occurrence.)
*------------------------------------------------------------------------*/
for (;;) {
term = softfloat_shortShiftLeft128( 0, sig64Z, 32 );
term = softfloat_add128( term.v64, term.v0, 0, (uint_fast64_t) q<<6 );
term = softfloat_mul128By32( term.v64, term.v0, q );
rem = softfloat_sub128( y.v64, y.v0, term.v64, term.v0 );
if ( ! (rem.v64 & UINT64_C( 0x8000000000000000 )) ) break;
--q;
}
qs[0] = q;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
q = (((rem.v64>>2) * recipSqrt32)>>32) + 2;
sigZExtra = (uint64_t) ((uint_fast64_t) q<<59);
term = softfloat_shortShiftLeft128( 0, qs[1], 53 );
sigZ =
softfloat_add128(
(uint_fast64_t) qs[2]<<18, ((uint_fast64_t) qs[0]<<24) + (q>>5),
term.v64, term.v0
);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( (q & 0xF) <= 2 ) {
q &= ~3;
sigZExtra = (uint64_t) ((uint_fast64_t) q<<59);
y = softfloat_shortShiftLeft128( sigZ.v64, sigZ.v0, 6 );
y.v0 |= sigZExtra>>58;
term = softfloat_sub128( y.v64, y.v0, 0, q );
y = softfloat_mul64ByShifted32To128( term.v0, q );
term = softfloat_mul64ByShifted32To128( term.v64, q );
term = softfloat_add128( term.v64, term.v0, 0, y.v64 );
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 20 );
term = softfloat_sub128( term.v64, term.v0, rem.v64, rem.v0 );
/*--------------------------------------------------------------------
| The concatenation of `term' and `y.v0' is now the negative remainder
| (3 words altogether).
*--------------------------------------------------------------------*/
if ( term.v64 & UINT64_C( 0x8000000000000000 ) ) {
sigZExtra |= 1;
} else {
if ( term.v64 | term.v0 | y.v0 ) {
if ( sigZExtra ) {
--sigZExtra;
} else {
sigZ = softfloat_sub128( sigZ.v64, sigZ.v0, 0, 1 );
sigZExtra = ~0;
}
}
}
}
return softfloat_roundPackToF128( state, 0, expZ, sigZ.v64, sigZ.v0, sigZExtra );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t f128_sub( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
#if ! defined INLINE_LEVEL || (INLINE_LEVEL < 2)
float128_t
(*magsFuncPtr)(
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
#endif
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
#if defined INLINE_LEVEL && (2 <= INLINE_LEVEL)
if ( signA == signB ) {
return softfloat_subMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
} else {
return softfloat_addMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
}
#else
magsFuncPtr =
(signA == signB) ? softfloat_subMagsF128 : softfloat_addMagsF128;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
#endif
}
+95
View File
@@ -0,0 +1,95 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float16_t f128_to_f16( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t frac64;
struct commonNaN commonNaN;
uint_fast16_t uiZ, frac16;
union ui16_f16 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
frac64 = fracF128UI64( uiA64 ) | (uiA0 != 0);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 ) {
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF16UI( &commonNaN );
} else {
uiZ = packToF16UI( sign, 0x1F, 0 );
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
frac16 = softfloat_shortShiftRightJam64( frac64, 34 );
if ( ! (exp | frac16) ) {
uiZ = packToF16UI( sign, 0, 0 );
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
exp -= 0x3FF1;
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x40 ) exp = -0x40;
}
return softfloat_roundPackToF16( sign, exp, frac16 | 0x4000 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+95
View File
@@ -0,0 +1,95 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float32_t f128_to_f32( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t frac64;
struct commonNaN commonNaN;
uint_fast32_t uiZ, frac32;
union ui32_f32 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
frac64 = fracF128UI64( uiA64 ) | (uiA0 != 0);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 ) {
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF32UI( &commonNaN );
} else {
uiZ = packToF32UI( sign, 0xFF, 0 );
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
frac32 = softfloat_shortShiftRightJam64( frac64, 18 );
if ( ! (exp | frac32) ) {
uiZ = packToF32UI( sign, 0, 0 );
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
exp -= 0x3F81;
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return softfloat_roundPackToF32( state, sign, exp, frac32 | 0x40000000 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+100
View File
@@ -0,0 +1,100 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float64_t f128_to_f64( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t frac64, frac0;
struct commonNaN commonNaN;
uint_fast64_t uiZ;
struct uint128 frac128;
union ui64_f64 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
frac64 = fracF128UI64( uiA64 );
frac0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 | frac0 ) {
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF64UI( &commonNaN );
} else {
uiZ = packToF64UI( sign, 0x7FF, 0 );
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
frac128 = softfloat_shortShiftLeft128( frac64, frac0, 14 );
frac64 = frac128.v64 | (frac128.v0 != 0);
if ( ! (exp | frac64) ) {
uiZ = packToF64UI( sign, 0, 0 );
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
exp -= 0x3C01;
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return
softfloat_roundPackToF64(
state, sign, exp, frac64 | UINT64_C( 0x4000000000000000 ) );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+85
View File
@@ -0,0 +1,85 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
int_fast32_t f128_to_i32( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64, sig0;
int_fast32_t shiftDist;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 );
sig0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
#if (i32_fromNaN != i32_fromPosOverflow) || (i32_fromNaN != i32_fromNegOverflow)
if ( (exp == 0x7FFF) && (sig64 | sig0) ) {
#if (i32_fromNaN == i32_fromPosOverflow)
sign = 0;
#elif (i32_fromNaN == i32_fromNegOverflow)
sign = 1;
#else
softfloat_raiseFlags( softfloat_flag_invalid );
return i32_fromNaN;
#endif
}
#endif
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
sig64 |= (sig0 != 0);
shiftDist = 0x4023 - exp;
if ( 0 < shiftDist ) sig64 = softfloat_shiftRightJam64( sig64, shiftDist );
return softfloat_roundToI32( state, sign, sig64, roundingMode, exact );
}
+95
View File
@@ -0,0 +1,95 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
int_fast64_t f128_to_i64( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64, sig0;
int_fast32_t shiftDist;
struct uint128 sig128;
struct uint64_extra sigExtra;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 );
sig0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
shiftDist = 0x402F - exp;
if ( shiftDist <= 0 ) {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( shiftDist < -15 ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig64 | sig0) ? i64_fromNaN
: sign ? i64_fromNegOverflow : i64_fromPosOverflow;
}
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
sig64 |= UINT64_C( 0x0001000000000000 );
if ( shiftDist ) {
sig128 = softfloat_shortShiftLeft128( sig64, sig0, -shiftDist );
sig64 = sig128.v64;
sig0 = sig128.v0;
}
} else {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
sigExtra = softfloat_shiftRightJam64Extra( sig64, sig0, shiftDist );
sig64 = sigExtra.v;
sig0 = sigExtra.extra;
}
return softfloat_roundToI64( state, sign, sig64, sig0, roundingMode, exact );
}
+86
View File
@@ -0,0 +1,86 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
uint_fast32_t
f128_to_ui32( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64;
int_fast32_t shiftDist;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 ) | (uiA0 != 0);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
#if (ui32_fromNaN != ui32_fromPosOverflow) || (ui32_fromNaN != ui32_fromNegOverflow)
if ( (exp == 0x7FFF) && sig64 ) {
#if (ui32_fromNaN == ui32_fromPosOverflow)
sign = 0;
#elif (ui32_fromNaN == ui32_fromNegOverflow)
sign = 1;
#else
softfloat_raiseFlags( softfloat_flag_invalid );
return ui32_fromNaN;
#endif
}
#endif
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
shiftDist = 0x4023 - exp;
if ( 0 < shiftDist ) {
sig64 = softfloat_shiftRightJam64( sig64, shiftDist );
}
return softfloat_roundToUI32( sign, sig64, roundingMode, exact );
}
+96
View File
@@ -0,0 +1,96 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
uint_fast64_t
f128_to_ui64( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64, sig0;
int_fast32_t shiftDist;
struct uint128 sig128;
struct uint64_extra sigExtra;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 );
sig0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
shiftDist = 0x402F - exp;
if ( shiftDist <= 0 ) {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( shiftDist < -15 ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig64 | sig0) ? ui64_fromNaN
: sign ? ui64_fromNegOverflow : ui64_fromPosOverflow;
}
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
sig64 |= UINT64_C( 0x0001000000000000 );
if ( shiftDist ) {
sig128 = softfloat_shortShiftLeft128( sig64, sig0, -shiftDist );
sig64 = sig128.v64;
sig0 = sig128.v0;
}
} else {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
sigExtra = softfloat_shiftRightJam64Extra( sig64, sig0, shiftDist );
sig64 = sigExtra.v;
sig0 = sigExtra.extra;
}
return softfloat_roundToUI64( state, sign, sig64, sig0, roundingMode, exact );
}
+96
View File
@@ -0,0 +1,96 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f32_to_f128( struct softfloat_state *state, float32_t a )
{
union ui32_f32 uA;
uint_fast32_t uiA;
bool sign;
int_fast16_t exp;
uint_fast32_t frac;
struct commonNaN commonNaN;
struct uint128 uiZ;
struct exp16_sig32 normExpSig;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA = uA.ui;
sign = signF32UI( uiA );
exp = expF32UI( uiA );
frac = fracF32UI( uiA );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0xFF ) {
if ( frac ) {
softfloat_f32UIToCommonNaN( state, uiA, &commonNaN );
uiZ = softfloat_commonNaNToF128UI( &commonNaN );
} else {
uiZ.v64 = packToF128UI64( sign, 0x7FFF, 0 );
uiZ.v0 = 0;
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! exp ) {
if ( ! frac ) {
uiZ.v64 = packToF128UI64( sign, 0, 0 );
uiZ.v0 = 0;
goto uiZ;
}
normExpSig = softfloat_normSubnormalF32Sig( frac );
exp = normExpSig.exp - 1;
frac = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ.v64 = packToF128UI64( sign, exp + 0x3F80, (uint_fast64_t) frac<<25 );
uiZ.v0 = 0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+64
View File
@@ -0,0 +1,64 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t i32_to_f128( int32_t a )
{
uint_fast64_t uiZ64;
bool sign;
uint_fast32_t absA;
int_fast8_t shiftDist;
union ui128_f128 uZ;
uiZ64 = 0;
if ( a ) {
sign = (a < 0);
absA = sign ? -(uint_fast32_t) a : (uint_fast32_t) a;
shiftDist = softfloat_countLeadingZeros32( absA ) + 17;
uiZ64 =
packToF128UI64(
sign, 0x402E - shiftDist, (uint_fast64_t) absA<<shiftDist );
}
uZ.ui.v64 = uiZ64;
uZ.ui.v0 = 0;
return uZ.f;
}
@@ -196,16 +196,20 @@ struct exp32_sig128
float128_t
softfloat_roundPackToF128(
struct softfloat_state *,
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast64_t );
float128_t
softfloat_normRoundPackToF128(
struct softfloat_state *,
bool, int_fast32_t, uint_fast64_t, uint_fast64_t );
float128_t
softfloat_addMagsF128(
struct softfloat_state *,
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
float128_t
softfloat_subMagsF128(
struct softfloat_state *,
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
float128_t
softfloat_mulAddF128(
File renamed without changes.
File renamed without changes.
+155
View File
@@ -0,0 +1,155 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
float128_t
softfloat_addMagsF128(
struct softfloat_state *state,
uint_fast64_t uiA64,
uint_fast64_t uiA0,
uint_fast64_t uiB64,
uint_fast64_t uiB0,
bool signZ
)
{
int_fast32_t expA;
struct uint128 sigA;
int_fast32_t expB;
struct uint128 sigB;
int_fast32_t expDiff;
struct uint128 uiZ, sigZ;
int_fast32_t expZ;
uint_fast64_t sigZExtra;
struct uint128_extra sig128Extra;
union ui128_f128 uZ;
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
expDiff = expA - expB;
if ( ! expDiff ) {
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 | sigB.v64 | sigB.v0 ) goto propagateNaN;
uiZ.v64 = uiA64;
uiZ.v0 = uiA0;
goto uiZ;
}
sigZ = softfloat_add128( sigA.v64, sigA.v0, sigB.v64, sigB.v0 );
if ( ! expA ) {
uiZ.v64 = packToF128UI64( signZ, 0, sigZ.v64 );
uiZ.v0 = sigZ.v0;
goto uiZ;
}
expZ = expA;
sigZ.v64 |= UINT64_C( 0x0002000000000000 );
sigZExtra = 0;
goto shiftRight1;
}
if ( expDiff < 0 ) {
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
uiZ.v64 = packToF128UI64( signZ, 0x7FFF, 0 );
uiZ.v0 = 0;
goto uiZ;
}
expZ = expB;
if ( expA ) {
sigA.v64 |= UINT64_C( 0x0001000000000000 );
} else {
++expDiff;
sigZExtra = 0;
if ( ! expDiff ) goto newlyAligned;
}
sig128Extra =
softfloat_shiftRightJam128Extra( sigA.v64, sigA.v0, 0, -expDiff );
sigA = sig128Extra.v;
sigZExtra = sig128Extra.extra;
} else {
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 ) goto propagateNaN;
uiZ.v64 = uiA64;
uiZ.v0 = uiA0;
goto uiZ;
}
expZ = expA;
if ( expB ) {
sigB.v64 |= UINT64_C( 0x0001000000000000 );
} else {
--expDiff;
sigZExtra = 0;
if ( ! expDiff ) goto newlyAligned;
}
sig128Extra =
softfloat_shiftRightJam128Extra( sigB.v64, sigB.v0, 0, expDiff );
sigB = sig128Extra.v;
sigZExtra = sig128Extra.extra;
}
newlyAligned:
sigZ =
softfloat_add128(
sigA.v64 | UINT64_C( 0x0001000000000000 ),
sigA.v0,
sigB.v64,
sigB.v0
);
--expZ;
if ( sigZ.v64 < UINT64_C( 0x0002000000000000 ) ) goto roundAndPack;
++expZ;
shiftRight1:
sig128Extra =
softfloat_shortShiftRightJam128Extra(
sigZ.v64, sigZ.v0, sigZExtra, 1 );
sigZ = sig128Extra.v;
sigZExtra = sig128Extra.extra;
roundAndPack:
return
softfloat_roundPackToF128( state, signZ, expZ, sigZ.v64, sigZ.v0, sigZExtra );
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
Loaded 100 of 663 files, more files were not shown because too many files have changed in this diff. Show more