Compare commits

...
375 Commits
Author SHA1 Message Date
Ryan Houdek ee0c1457d8 Docs: Update for release FEX-2310 2023-10-05 14:39:10 -07:00
Alyssa Rosenzweig 3413eb3d98 Merge pull request #3169 from Sonicadvance1/remove_constant_indirection
FEXCore: Support CpuState relative vector named constants
2023-10-05 08:22:48 -04:00
Ryan Houdek 2e0753a244 InstCountCI: Update for named vector constant optimization 2023-10-04 20:57:09 -07:00
Ryan Houdek 8a51bb7a61 FEXCore: Support CpuState relative vector named constants
The motivation towards just having a pointer array in CpuState was that
initialization was fairly cheap and that we have limited space inside
the encoding depending on what we want to do.

Initialization cost is still a concern but doing a memcpy of 128-bytes
isn't that big of a deal.

Limited space in CpuState, while a concern isn't a significant one.
   - Needs to currently be less than 1 page in size
   - Needs to be under the architectural offset limitations of loadstore
     scaled offsets. Which is 65KB for 128-bit vectors

Still keeps the pointer array around for cases when we would need
synthesize an address offset and it's just easier to load the
process-wide table.

The performance improvement here is removing the dependency in the
ldr+ldr chain. In microbenchmarks this has shown to have an improvement
of ~4% by removing this dependency chain on Cortex-X1C.
2023-10-04 20:56:29 -07:00
Ryan Houdek ee6debe8fd FEXCore: Adds DividePow2 helper 2023-10-04 20:56:29 -07:00
Mai 3ba1c7912c Merge pull request #3178 from Sonicadvance1/fix_avx_alias_precolour
Minor AVX optimizations
2023-10-04 21:31:20 -04:00
Ryan Houdek a408afaeb0 InstCountCI: Update for optimized AVX 2023-10-04 10:05:09 -07:00
Ryan Houdek fba7c4bedc IR/RA: Fixes register aliasing and pre-colouring for AVX
This is the cause of a bunch of redundant moves that shows up in
InstCountCI. Fixing this aliasing and pre-colouring issue causes a ton
of 256-bit operations to become optimal.
2023-10-04 10:04:06 -07:00
Ryan Houdek c52753e9c8 OpcodeDispatcher: Minor optimization in vzeroall
Using the cached zero value is less efficient than loading it in to the
register for all these cases.

Lets us use rename hardware more efficiently and removes a dependency
chain on a single register.

Original:
```
movi v2.2d, #0x0
mov z16.d, p7/m, z2.d
<... 16 more times>
mov z31.d, p7/m, z2.d
```

Result:
```
movi v16.2d, #0x0
<... 16 more times>
movi v31.2d, #0x0
```
2023-10-04 10:01:13 -07:00
Ryan Houdek e39634d314 Arm64: Fixes assert in VSQSHL/VSQSHR with SVE
When Dst != Vector then we need to pass Dst in to both Zd and Zdn.
Would have worked fine in a release build but assert build managed to
capture it.
2023-10-04 09:59:59 -07:00
Ryan Houdek 507cf82dad Merge pull request #3176 from neobrain/fix_thunks_unused_artifacts
Thunks: Only build guest target for libfex_thunk_test if FEXLinuxTests are enabled
2023-10-04 07:07:18 -07:00
Ryan Houdek 48fa4f1121 Merge pull request #3156 from neobrain/feature_thunk_data_layout_analysis
Thunks: Analyze data layout to detect platform differences
2023-10-04 07:06:49 -07:00
Tony Wasserka e06d609bf0 Thunks: Drop unused STRUCT_VERIFIER define from CMake 2023-10-03 11:43:29 +02:00
Tony Wasserka 0a09e04e33 Thunks: Only build guest target for libfex_thunk_test if FEXLinuxTests are enabled 2023-10-03 11:43:27 +02:00
Ryan Houdek a1a709f948 Merge pull request #3170 from Sonicadvance1/vixl_sim_instcountci
InstCountCI: Enable running on x86 hosts
2023-10-02 16:38:25 -07:00
Ryan Houdek 5925eef213 Github/InstCountCI: Enables x86 runner
To ensure we don't break this path for developers.
2023-10-02 16:26:14 -07:00
Ryan Houdek df369bd6a0 InstCountCI: Enable running on x86 hosts
This is a quality of life improvement for people that want to tinker
with the InstCountCI but they may not necessarily have an Arm64 device
available immediately for poking.

As long as the vixl disassembler is enabled then the InstCountCI tests
can run and get bit-accurate encodings just like on an Arm64 device.

This also ensures that behaviour is consistent with or without the vixl
simulator enabled which is very important when running on x86 hosts.
2023-10-02 16:26:14 -07:00
Ryan Houdek 978489fce1 InstCountCI: Explicitly disable SVE256 for one test group
These instructions are specifically testing the SVE128 implementations,
don't want SVE256 mucking up the instructions.
2023-10-02 16:26:14 -07:00
Ryan Houdek d5a4d9b17f InstCountCI: Adds option to disable cssc for tests
One x87 instruction was using CSSC abs
2023-10-02 16:26:14 -07:00
Ryan Houdek 9933ef07ea Tools: Enable indirect vixl runtime calls if simulator is used
So tests can still run.
2023-10-02 16:26:14 -07:00
Ryan Houdek 6964e65660 HostFeatures: Hardcode icache and dcache line size on x86
64-byte is effectively part of x86's ABI anyway. No need to query it for
our uses.
2023-10-02 16:26:14 -07:00
Ryan Houdek 11db8e7506 FEXCore: Wire up the new option to disable vixl indirect runtimes
Also so it compiles without the vixl simulator enabled.
2023-10-02 16:26:12 -07:00
Ryan Houdek b6b5e93dbb Config: Adds an option to disable vixl sim indirect runtime calls 2023-10-02 16:23:11 -07:00
Ryan Houdek 935b3a313a Merge pull request #3171 from Sonicadvance1/merge_dispatcher
FEXCore: Merge Arm64Dispatcher in to Dispatcher
2023-10-02 16:22:36 -07:00
Tony Wasserka fe681ab335 unittests/ThunkLibs: Specify clang resource directory when compiling test code 2023-10-02 22:18:23 +02:00
Tony Wasserka 2d9e816ff5 unittests/ThunkLibs: Add various tests for structs repacking and for void parameters 2023-10-02 22:18:23 +02:00
Tony Wasserka b04b0549a9 unittests/ThunkLibs: Add data layout tests 2023-10-02 22:18:22 +02:00
Tony Wasserka 2b472cb962 Thunks/gen: Enforce type compatibility for function parameters 2023-10-02 22:18:22 +02:00
Tony Wasserka 7f931b5623 Thunks/gen: Add detection logic for data layout differences
This runs the data layout analysis pass added in the previous change twice:
Once for the host architecture and once for the guest architecture. This
allows the new DataLayoutCompareAction to query architecture differences for
each type, which can then be used to instruct code generation accordingly.

Currently, type compatibility is classified into 3 categories:
* Fully compatible (same size/alignment for the type itself and any members)
* Repackable (incompatibility can be resolved with emission of automatable
  repacking code, e.g. when struct members are located at differing offsets
  due to padding bytes)
* Incompatible
2023-10-02 22:18:22 +02:00
Tony Wasserka 070fa9f924 Thunks/gen: Add data layout analysis
This adds a ComputeDataLayout function that maps a set of clang::Types
to an internal representation of their data layout (size, member list, ...).
2023-10-02 22:18:22 +02:00
Tony Wasserka 371bf50c76 Thunks/gen: Track data types passed across architecture boundaries
The set of these types is tracked in AnalysisAction, to which extensive
verification logic is added to detect potential incompatibilities and to
enforce use of annotatations where needed.
2023-10-02 22:18:22 +02:00
Tony Wasserka d65d29903b Thunks/gen: Rename EmitOutput to OnAnalysisComplete 2023-10-02 22:03:10 +02:00
Tony Wasserka 7791e0090d Thunks: Disable 32-bit host thunks
These are not supported yet.
2023-10-02 22:03:10 +02:00
Alyssa Rosenzweig 02da6d6ce7 Merge pull request #3174 from Sonicadvance1/remove_steam_appconfig
AppConfig: Removes Steam config
2023-10-01 18:48:30 -04:00
Ryan Houdek a478cbb694 AppConfig: Removes Steam config
This was only required on x86 devices trying to escape the emulation.
Since x86 is now remove, this is entirely unnecessary.

When Steam launches applications with `/bin/sh`, this will remain under
the emulation and not escape these days.
2023-10-01 08:46:53 -07:00
Ryan Houdek 3a25dd6d2b Merge pull request #3173 from CallumDev/x87f64-fabs
X87F64: Implement FABS with vector instruction
2023-10-01 01:54:11 -07:00
CallumDev 9c25db83d9 JIT: VectorOps remove extraneous element size logs 2023-10-01 15:03:21 +10:30
CallumDev 7346476546 Update InstCountCI 2023-10-01 14:41:13 +10:30
CallumDev c42b581378 X87F64: Implement FABS with vector instruction 2023-10-01 14:39:55 +10:30
Ryan Houdek ccfd770d9d Merge pull request #3172 from CallumDev/x87f64-opts
X87F64: Use Bfe for rounding mode, FCHS use float instruction
2023-09-30 18:41:29 -07:00
CallumDev d4a623a3fb InstCountCI Update 2023-10-01 11:22:18 +10:30
CallumDev c09c25005e X87F64: Use Bfe for rounding mode, FCHS use float instruction 2023-10-01 11:11:33 +10:30
Ryan Houdek 90570fd5f4 FEXCore: Merge Arm64Dispatcher in to Dispatcher
With the removal of the x86 JIT, there is no need to have these be
independent classes.

Merges the Arm64Dispatcher in to the base Dispatcher class.
No functional change, just moving code.
2023-09-30 09:31:55 -07:00
Mai ab4642af38 Merge pull request #3167 from Sonicadvance1/gatherqdps
unittests/ASM: Implements tests for vpgatherqd/vgatherqps
2023-09-29 12:16:43 -04:00
Mai d94e5ce7f4 Merge pull request #3168 from Sonicadvance1/gatherqqpd
unittests/ASM: Implements tests for vpgatherqq/vgatherqpd
2023-09-29 12:16:12 -04:00
Mai dad7086fd0 Merge pull request #3166 from Sonicadvance1/gatherdqpd
unittests/ASM: Implements tests for vpgatherdq/vgatherpq
2023-09-29 12:15:39 -04:00
Ryan Houdek a21def7d74 unittests/ASM: Implements tests for vpgatherqq/vgatherqpd
Similar to previous tests, vpgatherqq and vgatherqpd are equivalent
instructions. So the tests are the same with the mnemonic changed.

This adds tests for an additional two sets of instructions. Getting us
full coverage of all eight instructions if we include the tests from
PR #3167 and #3166

Tests the same things as described in #3165

In addition, since these tests use 64-bit indices for address
calculation, we can easily generate and indice vector that tests
overflow. So every test at every displacement ALSO gains an additional
overflow test to ensure correct behaviour around pointer overflow
calculation.
2023-09-29 08:04:47 -07:00
Ryan Houdek 0d8d5444a4 unittests/ASM: Implements tests for vpgatherqd/vgatherqps
Similar to previous tests, vgatherqd and vgatherqps are equivalent
instructions. So the tests are the same with the mnemonic changed.

This adds tests for an additional two sets of instructions, Getting us
up to six total over the eight if we include the tests from #3166.

Tests the same things as described in #3165

In addition, since these tests use 64-bit indices for address
calculation, we can easily generate and indice vector that tests
overflow. So every test at every displacement ALSO gains and additional
overflow test to ensure correct behaviour around pointer overflow
calculation.
2023-09-29 07:20:07 -07:00
Ryan Houdek eedfad5036 unittests/ASM: Implements tests for vpgatherdq/vgatherpq
Just like the previous tests, vpgatherdq and vgatherpq are equivalent
instructions. So the tests are the same except for the instruction
mnemonic again.

This adds unittests for two more of the eight gather instructions.
Getting us up to testing four in total.
Specifically this adds tests for 32-bit indices while loading 64-bit
element instructions.

Same thing as PR #3165 for what it tests versus doesn't.
2023-09-28 22:49:03 -07:00
Ryan Houdek 85da0f0640 Merge pull request #3165 from Sonicadvance1/gatherddps
unittests/ASM: Implements tests for vpgatherdd/vgatherps
2023-09-28 22:44:38 -07:00
Ryan Houdek 9a01b440e3 unittests/ASM: Implements tests for vpgatherdd/vgatherps
vpgatherdd and vgatherps are effectively the same instructions, so the
tests are the same except for the instruction mnemonic.

This adds unit tests for two of the eight gather instructions.
Specifically this adds tests for the 32-bit indices loading 32-bit
elements instructions.

What it tests:
- Tests all displacement scales
- Tests multiple mask arrangements
- Ensures the mask register is zero'd after the instruction

What it doesn't test:
- Doesn't test address size calculation overflow
   - Only would happen on 32-bit with 32-bit indices, or /really/ high
     base addresses
   - The instruction should behave as a mask to the address size
   - Effectively behaves like `(uint64_t)(base + index << ilog2(scale))`
   - Better idea is to just not expose AVX to 32-bit applications
- Doesn't test VSIB immediate displacement
   - This just ends up being base_addr + imm so it isn't too interesting
   - We can add more tests in the future if we think we messed that up
- Doesn't test partial fault behaviour
   - Because that's a nightmare.

Specifically keeps each instruction test small and isolated so if a
single register fails it is very easily to nail down which operation did
it.
I know some of our ASM tests do a chunk of work and spit out a result at
the end which can be difficult to debug in some cases. Didn't want to do
that which is why the tests are spread out across 16 files for these
single class of instructions.
2023-09-28 19:58:34 -07:00
Ryan Houdek 228ee7fa47 TestHarnessRunner: Support AVX2 flag detection 2023-09-28 19:58:34 -07:00
Ryan Houdek 98789a8039 FEXCore: Implement support for AVX2 feature detection 2023-09-28 19:57:08 -07:00
Ryan Houdek 14398742c3 Merge pull request #3164 from neobrain/fix_thunks_asan
Thunks: Fix AddressSanitizer build
2023-09-28 12:05:55 -07:00
Tony Wasserka 5a7e3192da Thunks: Fix AddressSanitizer build 2023-09-28 15:13:03 +02:00
Ryan Houdek 6b4ff4ae81 Merge pull request #3163 from alyssarosenzweig/opt/ascii-flags
Optimize ASCII flags
2023-09-27 10:42:47 -07:00
Ryan Houdek d1d3de80d1 Merge pull request #3157 from alyssarosenzweig/opt/unmask-in
OpcodeDispatcher: Don't mask logic op inputs
2023-09-27 10:38:12 -07:00
Alyssa Rosenzweig 2e32e1367d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:55:57 -04:00
Alyssa Rosenzweig 711583aa76 OpcodeDispatcher: Optimize PTEST flags
Zero NZCV first to avoid RMW.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:55:57 -04:00
Alyssa Rosenzweig 3efac9646c OpcodeDispatcher: Optimize ASCII flags
Make the zeroing of undefined NZCV more obvious. Mitigates regressions from
future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:31:31 -04:00
Alyssa Rosenzweig 095a362046 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:30:09 -04:00
Alyssa Rosenzweig 3bb64c64e3 OpcodeDispatcher: Don't mask for TEST
Like AND.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:30:02 -04:00
Alyssa Rosenzweig a4de164944 OpcodeDispatcher: Use lshr for ah/bh with AllowUpperGarbage
If we ever get around to fusing ops with shifts in the ConstProp optimizer (may
or may not be worthwhile), this will delete an instruction from things like "or
al, bh".

Even though lsr is the same speed as bfe on Firestorm, I feel if you ask for
garbage you should get garbage C:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:28:01 -04:00
Alyssa Rosenzweig 45a645fbbc OpcodeDispatcher: Don't mask logic op inputs
Pointless, upper bits ignored anyway. Deletes piles of uxt and even some 32-bit
instruction moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:12:22 -04:00
Alyssa Rosenzweig 92211bf8c6 OpcodeDispatcher: Add AllowUpperGarbage option
To load 8-bit sources without bfe'ing for al/bl/cl if the caller knows it
doesn't need masking behaviour, but without lying about the size so the extract
for ah/bh/ch will still work properly.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:08:20 -04:00
Alyssa Rosenzweig 728d3f8ac7 InstCountCI: Add a case with a hi 8-bit reg
Noticeably different code pattern.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 18:33:55 -04:00
Ryan Houdek ca87d8688d Merge pull request #3153 from alyssarosenzweig/opt/adcs
Use adcs
2023-09-26 09:57:01 -07:00
Ryan Houdek e32601f49d Merge pull request #3161 from neobrain/fix_ctest_silent_failures
unittests: Instruct CTest to print output from tests on failure
2023-09-26 08:26:15 -07:00
Tony Wasserka f4dd456c80 unittests: Instruct CTest to print output from tests on failure 2023-09-26 17:16:28 +02:00
Alyssa Rosenzweig 7b22dbfe24 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 10:05:59 -04:00
Alyssa Rosenzweig 7a06cc9727 IR: Use adcs/sbcs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 09:06:46 -04:00
Ryan Houdek 8b3881b5db Merge pull request #3154 from alyssarosenzweig/opt/smol-carry
Optimize 8/16-bit CF calculation
2023-09-26 05:49:07 -07:00
Ryan Houdek 76d4637d9c Merge pull request #3159 from neobrain/feature_update_vulkan
Thunks: Update Vulkan thunk to v1.3.261.1
2023-09-26 05:20:18 -07:00
Alyssa Rosenzweig 0d12cce74f Merge pull request #3158 from Sonicadvance1/unittest_for_3153
unittests/ASM: Adds unit test caught by #3153
2023-09-26 08:15:40 -04:00
Tony Wasserka 04592af609 Thunks: Update Vulkan thunk to v1.3.261.1 2023-09-26 12:14:58 +02:00
Ryan Houdek d8366c04dc unittests/ASM: Adds unit test caught by #3153 2023-09-26 00:28:45 -07:00
Ryan Houdek 533f35934c Merge pull request #3155 from neobrain/opt_thunks_rebuilds
Thunks: Avoid recompiling thunk interfaces on FEXLoader changes
2023-09-25 19:21:09 -07:00
Alyssa Rosenzweig 35bb7cc801 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 19:41:31 -04:00
Alyssa Rosenzweig 5facb21d30 OpcodeDispatcher: Don't mask small add/sub carries
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.

Deletes a pile of code both in FEX and the output :-)

ADC/SBC could probably get similar treatment later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 18:25:30 -04:00
Tony Wasserka adead832a5 Thunks: Avoid recompiling thunk interfaces on FEXLoader changes
The interface files themselves don't use FEXLoader. Only the final library
does.
2023-09-25 23:04:09 +02:00
Ryan Houdek 5eed24a242 Merge pull request #3152 from Sonicadvance1/instcountci_x87_f64
InstCountCI: Support f64 reduced precision mode tests
2023-09-24 19:29:37 -07:00
Ryan Houdek 7907f70ed2 InstCountCI: Adds new x87 reduced precision mode tests 2023-09-24 18:50:05 -07:00
Ryan Houdek 7141332f6f InstCountCI: Support setting environment variables in tests
This will allow us to enable FEX options through environment variables
just like the ASM tests.
2023-09-24 18:50:01 -07:00
Ryan Houdek 234e029391 Merge pull request #3145 from Sonicadvance1/optimize_inline_calls
PassManager: Optimize out CPUID and XGetBV calls
2023-09-24 18:09:18 -07:00
Ryan Houdek 19a7b514e6 Merge pull request #3150 from alyssarosenzweig/opt/ornror
Optimize PF calculation in lahf
2023-09-24 18:05:57 -07:00
Ryan Houdek 220761a0e8 Merge pull request #3151 from Sonicadvance1/unique_name_workflow_jobs
Github: Changes jobs to have unique names
2023-09-24 18:03:57 -07:00
Alyssa Rosenzweig cbd4daddff InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig c8519b0b87 OpcodeDispatcher: Remove LoadPF
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig 68d32ad70d OpcodeDispatcher: Optimize PF in lahf
Use the raw popcount rather than the final PF and use some sneaky bit math to
come out 1 instruction ahead.

Closes #3117

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Ryan Houdek 62890f148f Github: Changes jobs to have unique names
These overlapping names make it impossible to ensure all checks are
required to pass before merge.

Unique names will fix this.
2023-09-24 17:52:47 -07:00
Alyssa Rosenzweig 1f02a6da34 IR: Add Ornror op
Mostly copypaste of Orlshl... we really should deduplicate this mess somehow.
Maybe a shift enum on the core Or op?

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:47:50 -04:00
Alyssa Rosenzweig 86063411dc Revert "OpcodeDispatcher: Use plain Lshl for flags"
This logic is unused since 8adfaa9aa ("OpcodeDispatcher: Use SelectCC for x87"),
which addressed the underlying issue.

This reverts commit df3833edbe.
2023-09-24 20:47:50 -04:00
Ryan Houdek 9968e6431f Passes: Rename SyscallOptimization
This is now inlining multiple external calls out of the JIT. Rename it
to InlineCallOptimization.
2023-09-24 17:25:38 -07:00
Ryan Houdek ff24f64b2a PassManager: Optimize out CPUID and XGetBV calls
If we const-prop the required functions and leafs then we can directly
encode the CPUID information rather than jumping out of the JIT.
In testing almost all CPUID executions const-prop which function is
getting called. Worst case that I found was only 85% const-prop rate.

This isn't quite 100% optimal since we need to call the RCLSE and
Constprop passes after we optimize these, which would remove some
redundant moves.

Sadly there seems to be a bug in the constprop pass that starts crashing
applications if that is done.
Easily enough tested by running Half-Life 2 and it immediately hitting
SIGILL.

Even without this optimization, this is stil a significant savings since
we aren't jumping out of the JIT anymore for these optimized CPUIDs.
2023-09-24 17:25:38 -07:00
Ryan Houdek e9a7ef2534 CPUID: Describe CPUID functions if they return constant state or not
Most CPUID routines return constant data, there are four that don't.
Some CPUID functions also need the leaf descriptor, so we need to
describe that as well.

Functions that don't return constant data:
- function 1Ah - Returns different data depending on current CPU core
- function 8000_000{2,3,4} - Different data based on CPU core

Functions that need leaf constprop:
- 4h, 7h, Dh, 4000_0001h, 8000_001Dh
2023-09-24 17:25:38 -07:00
Ryan Houdek 842c57e221 CPUID: Constify some functions
These don't modify CPUIDEmu state.
2023-09-24 17:25:38 -07:00
Ryan Houdek 93aeb157b4 Merge pull request #3149 from Sonicadvance1/fail_on_change
InstCountCI: Fail CI if there was any difference.
2023-09-24 17:23:52 -07:00
Ryan Houdek 02ff9f200c InstCountCI: Upload diff and check for failure 2023-09-24 17:14:08 -07:00
Ryan Houdek f65b40f298 InstCountCI: Fail if inst count has changed 2023-09-24 17:14:08 -07:00
Ryan Houdek c38beff826 Merge pull request #3148 from Sonicadvance1/add_negative_primaries
InstCountCI: Adds negative immediate primary tests
2023-09-24 17:13:34 -07:00
Ryan Houdek 94c22b2269 InstCountCI: Adds negative immediate primary tests
Noticed these were missing
2023-09-24 17:02:58 -07:00
Ryan Houdek bee97309f6 Merge pull request #3147 from alyssarosenzweig/opt/0924
More opts to the dispatcher + 1 to the JIT
2023-09-24 17:01:37 -07:00
Alyssa Rosenzweig 331941dec6 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig 8798e0cba0 Arm64: Rewrite Set/GetRoundingMode
I went auditing for places to use cset and what I found was hot garbage.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig c5fc03dac4 OpcodeDispatcher: Use cset for blsr/etc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig e63871ed2e OpcodeDispatcher: Handle sub in CalculateOF
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig ea8b7633eb OpcodeDispatcher: Optimize OF calc of immediates
If we know the sign of one of the sources, we can do better when calculating OF.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 18:16:09 -04:00
Ryan Houdek e795ec683d Merge pull request #3139 from Sonicadvance1/workaround
FEXServerClient: Adds back ServerSocketPath config option
2023-09-23 17:09:04 -07:00
Ryan Houdek 6dc5c0d3be Merge pull request #3144 from Sonicadvance1/optimize_redundant_store_load
RCLSE: Optimize redundant store->load operations
2023-09-23 17:06:10 -07:00
Ryan Houdek eb5e0be569 FEXServerClient: Adds back ServerSocketPath config option
This option was disabled a few months ago when we switched the server
socket from a filesystem unix socket to an abstract socket.
This partially broke our chroot scripts which relied on this option
existing.

Readds support for an explicitly named abstract socket named from
config.

This is a workaround for dealing with chroots that change users.
They end up changing a user while doing operations and then can't
connect to the FEXServer anymore because environment variables have been
wiped away.
2023-09-23 16:59:58 -07:00
Ryan Houdek be3ff804a6 InstCountCI: Update for optimization 2023-09-23 06:11:35 -07:00
Ryan Houdek 9ab2967d71 Arm64: Fixes wide shifts
movprfx is invalid to use when the source register matches the movprfx
destination.

This was getting picked up on by `TwoByte/0F_D1.asm` now that RCLSE is
working better now.
2023-09-23 06:06:18 -07:00
Ryan Houdek d01b457727 RCLSE: Optimize redundant store->load operations
The bug that was causing crashes with this was due to inline syscalls.
Now that this is fixed we can re-enable store->load operations.

This allows constant propagation to work significantly better, which
means inline syscalls start working again. This can significantly
improve syscall performance in some cases.

This is most likely to improve performance in dxsetup and vc_redist but
hard to get a real profile.

Additionally this will let us inline cpuid results in the future which
is pretty nice.
2023-09-23 06:06:18 -07:00
Mai 4e9a114858 Merge pull request #3142 from Sonicadvance1/inline_syscall_fix
Arm64: Fixes inline syscalls
2023-09-23 09:03:49 -04:00
Mai 72d092e951 Merge pull request #3141 from Sonicadvance1/fix_simm9_range
ConstProp: Fixes unscaled signed 9-bit range
2023-09-23 09:03:01 -04:00
Mai da3e172857 Merge pull request #3140 from Sonicadvance1/fix_core_sanitization
Config: Fixes core sanitization
2023-09-23 09:01:42 -04:00
Ryan Houdek 28fa0bda31 Arm64: Fixes inline syscalls
Ever since we reordered registers in `X86Enums.h` this has silently been
broken. This wasn't hit because RCLSE has been broken ever since SRA was
added, so inlinesyscalls just weren't ever happening.

Quick fix while I think of a way to more strictly correlate these
registers so it doesn't happen again.
2023-09-23 02:56:32 -07:00
Ryan Houdek 1f2a3cfa8b ConstProp: Fixes unscaled signed 9-bit range
The range was slightly incorrect which mostly wouldn't have caused
issues.

The lowest byte would have just generated slightly less optimal code.
The upper byte could have generated broken code, which our CI couldn't
catch since TSO instructions only get enabled when multiple threads are
in-flight.

Easy enough to fix.
2023-09-23 01:13:54 -07:00
Ryan Houdek 571b0fe47e Config: Fixes core sanitization
This would have caused core to try and initialize a custom core on
Arm64, which causes a std::function assert because it doesn't support
that.

Users would likely get hit by this immediately since we deleted the
interpreter and shifted all the core numbers.
2023-09-23 00:52:23 -07:00
Ryan Houdek 86ad35c418 Merge pull request #3138 from alyssarosenzweig/opt/train
Requiem for the x86 jit
2023-09-22 16:33:15 -07:00
Alyssa Rosenzweig 0b27029c3f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:10:41 -04:00
Alyssa Rosenzweig 223a6562ff IR: Support <32-bit TestNZ
Originally this was going to use setf8/setf16, but it looks like the approach of
shift-and-test turns out to be faster. As a bonus this is a nice delete-the-code
win :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:08:26 -04:00
Alyssa Rosenzweig b1231c24ef OpcodeDispatcher: Omit AF xor for common constants
The only reason we need to XOR arguments for AF is to get bit 4 correct. But if
the operand in question is known to have bit 4 clear, the XOR will be an
effective no-op and can be skipped. This saves an instruction in a bunch of
common cases, like inc/dec. If we dedicated a register to AF to eliminate the
store, we would not save an instruction from this but would still come out ahead
due to an eor turning into a (zero cycle?) mov that can be handled by the
renamer.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:08:26 -04:00
Alyssa Rosenzweig 699aa85c4b OpcodeDispatcher: Opt PF selection
Fold the and in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Alyssa Rosenzweig 2d65a3677b OpcodeDispatcher: Optimize NZCV selects
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Alyssa Rosenzweig 2a2619c0f5 IR: Add bit masking selects
Add new synthetic condition codes that do an AND as their relational operator,
testing the result. This is 1 IR op for things like

  (A & B) == 0 ? C : D

This can translate to

  tst A, B
  csel A, B, eq

In the future, if A is the NZCV register and B is a supported immediate, eg

  (NZCV & 0x80000000) == 0 ? C : D

this will be able to translate to a single instruction with the appropriate
condition

  csel A, B, pl

but that needs RA support.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Ryan Houdek 797c890ff6 Merge pull request #2874 from bylaws/wowfex
Add WOW64 JIT frontend
2023-09-22 15:47:59 -07:00
Ryan Houdek 879b41c184 Merge pull request #3134 from Sonicadvance1/remove_x86_jit
FEXCore: Removes x86 JIT.
2023-09-22 15:36:47 -07:00
Ryan Houdek 0fbf403787 Adds back in host testharnessrunner CI
Necessary for asm tests to still run in the host "core".
Useful for ensuring correct behaviour of our assembly tests.
2023-09-22 14:46:03 -07:00
Billy Laws 04cf418452 Windows: Add SPDX license identifiers 2023-09-22 10:12:40 -07:00
Billy Laws 057a7c6ee8 WOW64: Implement thread suspension handling
This provides more robust handling than a signal based approach, as the
suspender is able to wait for the suspendee to reach a suitable position and
flush its context to memory before returning.
2023-09-22 10:12:40 -07:00
Billy Laws 3d6955592b WOW64: Implement partial self-modifying code handling
This should support most simple cases of SMC, however programs which make use
of separate shared memory mappings for writing and execution are not handled.
The overall approach is the same as is done for linux, where RWX mappings are
protected to RX and then when a write occurs the signal handler invalidates the
faulting page and reprotects it to RWX until code in that page is jitted again.
2023-09-22 10:12:40 -07:00
Billy Laws f57aee0a62 WOW64: Add a templated interval list implementation
Stores binary intervals in a sorted vector container, to be used for SMC
handling.
2023-09-22 10:12:40 -07:00
Billy Laws c978fdd12f WOW64: Implement basic code invalidation handling 2023-09-22 10:12:40 -07:00
Billy Laws 19713bd20a WOW64: Implement exception handling with context restoration
When an exception occurs, pretend that we were just at the point of JIT entry
so the stack can be unwound to the wow64 SEH handler, which then handles
dispatching the exception to the x86 guest with the restored context.
2023-09-22 10:12:40 -07:00
Billy Laws 22b1fea96d WOW64: Handle unaligned atomic accesses
This is done in EnsureConsistentState rather than as a VEH to avoid needing to
go through all of wine's exception handling logic for such a hot path.
2023-09-22 10:12:40 -07:00
Billy Laws be4fcaf65c WOW64: Report CPU features based off of the emulated cpuid 2023-09-22 10:12:40 -07:00
Billy Laws 2add8a7751 Windows: Introduce a barebones FEXCore-based WOW64 BT module
This allows for running x86 applications under wine without having to run all
of wine under FEX. The JIT is invoked when running application code and then
left when handling NT syscalls or unix calls to e.g. the Vulkan driver.
2023-09-22 10:12:40 -07:00
Billy Laws 9612133088 Windows: Generate import libraries for private ntdll and wow64 APIs
The MinGW supplied import libraries are incomplete and miss a lot of
functions necessary to implement lower level windows code. To avoid
needing to many resolve every function, pull in .def files from wine
that detail the entire ntdll and wow64 APIs.
2023-09-22 10:12:40 -07:00
Billy Laws f46fd42977 Windows: Add a minimal set of wine-derived headers
These are cut down versions of wine headers containing only what is necessary
for WOW. This shouldn't carry any license implications for FEX, as per the
LGPLv3 license:

```
The object code form of an Application may incorporate material from a header
file that is part of the Library. You may convey such object code under terms
of your choice, provided that, if the incorporated material is not limited to
numerical parameters, data structure layouts and accessors, or small macros,
inline functions and templates (ten or fewer lines in length), you do both of
the following:

a) Give prominent notice with each copy of the object code that the Library is
used in it and that the Library and its use are covered by this License.
b) Accompany the object code with a copy of the GNU GPL and this license
document.
```
2023-09-22 10:12:40 -07:00
Billy Laws 51f8c83c76 Context: Add an alternative thread-oriented execute function 2023-09-22 10:12:40 -07:00
Billy Laws d641d3f61e OpcodeDispatcher: Avoid redundantly passing args to WIN32 ABI syscalls 2023-09-22 10:12:39 -07:00
Ryan Houdek 02ae59a348 github: Disables default build test on x64 2023-09-21 18:30:03 -07:00
Ryan Houdek 64df9e31c6 github: Remove mingw tests from x86 CI 2023-09-21 18:30:03 -07:00
Ryan Houdek d32bb993a8 github: Remove glibc fault tests from x86 CI 2023-09-21 18:30:03 -07:00
Ryan Houdek b5cc9a12f2 FEXCore: Removes x86 JIT.
This is blocking performance improvements. This backend is almost
unilaterally unused except for when I'm testing if games run on Radeon
video drivers.

Hopefully AmpereOne and Orin/Grace can fulfill this role when they
launch next year.
2023-09-21 18:30:02 -07:00
Ryan Houdek 65b6df9dbb Merge pull request #3133 from Sonicadvance1/remove_vestigial_interpreter
FEXCore: Removes vestigial Interpreter code
2023-09-21 18:15:32 -07:00
Ryan Houdek 31564354b1 FEXCore: Removes vestigial Interpreter code 2023-09-21 15:49:49 -07:00
Ryan Houdek fea72ce19c Merge pull request #3120 from Sonicadvance1/more_optimal_x87
FEXCore: Support preserve_all ABI for interpreter fallbacks
2023-09-21 15:35:37 -07:00
Ryan Houdek 2b7e1d10ec Merge pull request #3131 from Sonicadvance1/optimize_btr
OpcodeDispatcher: Optimize lock btr
2023-09-21 15:06:55 -07:00
Ryan Houdek 5444810d64 Merge pull request #3132 from alyssarosenzweig/opt/orlshl
Optimize reconstructing x87, harder
2023-09-21 15:02:37 -07:00
Ryan Houdek 4a2ceabfdd InstCountCI: Add atomic bit test instructions
These all can likely be more optimal.
2023-09-21 14:54:51 -07:00
Ryan Houdek 1a4d1d820b OpcodeDispatcher: Optimize lock btr
This is an atomicFetchCLR, removes two mvn instructions that are back to
back negating the source.

We didn't have this instruction combination in InstCountCI so will be a
bit hard to see.
2023-09-21 14:54:51 -07:00
Ryan Houdek 0ae4bbb9c5 IR: Implements support for AtomicFetchCLR
This is the native ARM operation rather than fetchAnd. Will make an
instruction an instruction slightly more optimal.
2023-09-21 14:54:51 -07:00
Ryan Houdek 7d99eb05c6 Merge pull request #3128 from alyssarosenzweig/rm/interp
FEXCore: Gut interpreter
2023-09-21 14:51:44 -07:00
Alyssa Rosenzweig 8247ded2cf unittests: Remove stale comments
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 12:48:12 -04:00
Alyssa Rosenzweig c52741c813 FEXCore: Gut interpreter
It is scarcely used today, and like the x86 jit, it is a significant
maintainence burden complicating work on FEXCore and arm64 optimization. Remove
it, bringing us down to 2 backends.

1 down, 1 to go.

Some interpreter scaffolding remains for x87 fallbacks. That is not a problem
here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 12:48:12 -04:00
Alyssa Rosenzweig 75ffbc16f2 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:14:08 -04:00
Alyssa Rosenzweig 1596e33f58 OpcodeDispatcher: Remove pointless or
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig 07d03f1610 OpcodeDispatcher: Don't opencode bfe, badly
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig a8b48dcacd OpcodeDispatcher: Swap some selects
...if it lets us use cset.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig bb87b2a19d OpcodeDispatcher: Use more Orlshl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig 19eff62c77 OpcodeDispatcher: Use orlshl for FCW
Potentially easier on the RA (bfi has a tied operand), mostly whatever here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 08:55:25 -04:00
Mai 5fc8699db9 Merge pull request #3130 from Sonicadvance1/optimize_fsw
OpcodeDispatcher: Optimize reconstructing FSW
2023-09-21 08:35:16 -04:00
Mai 43fd159689 Merge pull request #3129 from Sonicadvance1/remove_non_explicit_selectcc
OpcodeDispatcher: Removes non-explicit SelectCC function
2023-09-21 08:33:30 -04:00
Ryan Houdek 758820ca86 InstCountCI: Update for optimized FSW reconstruction 2023-09-21 02:27:04 -07:00
Ryan Houdek 5664195e49 OpcodeDispatcher: Optimize reconstructing FSW
Minor optimization using Bfi to insert C0, C1, C2, & C3
2023-09-21 02:07:27 -07:00
Ryan Houdek 683daefc15 InstCountCI: Minor changes 2023-09-21 01:57:08 -07:00
Ryan Houdek 8e9e87f631 OpcodeDispatcher: Removes non-explicit SelectCC function
Renames the explicit sized one to `SelectCC`
Cleans up a bit of duplicated code.
2023-09-21 01:56:38 -07:00
Ryan Houdek 0a0865eb1c InstCountCI: Update for minor change 2023-09-20 18:51:18 -07:00
Ryan Houdek d588d41ab9 InterpreterFallbacks: Converts X87 and String ops to preserve_all
This improves performance!
2023-09-20 18:51:18 -07:00
Ryan Houdek 8aa8d597f6 Arm64: Supports jumping out of the JIT with preserve_all ABI
This improves perferformance when jumping out of the Arm64 JIT by
reducing the number of registers we need to save.
2023-09-20 18:51:18 -07:00
Ryan Houdek 67680d71a4 Merge pull request #3125 from Sonicadvance1/spdx_fexcore
FEXCore: Adds SPDX identifier
2023-09-19 17:42:07 -07:00
Ryan Houdek d86f41e29a Merge pull request #3124 from Sonicadvance1/spdx_fexcore_include
FEXCore/Include: Adds SPDX identifier
2023-09-19 17:41:59 -07:00
Ryan Houdek ba56e514bd Merge pull request #3123 from Sonicadvance1/spdx_fex_linux
FEX: Moves Linux utils and adds spdx
2023-09-19 17:41:51 -07:00
Ryan Houdek 9f5f09b772 Merge pull request #3122 from Sonicadvance1/spdx_fex_common
FEX/Common: Adds SPDX identifier
2023-09-19 17:41:44 -07:00
Ryan Houdek ddf4b5cbd4 Merge pull request #3121 from Sonicadvance1/spdx_tools
FEX/Tools: Adds SPDX identifier
2023-09-19 17:41:34 -07:00
Ryan Houdek e4613477b1 FEXCore/Interface/Core: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek d18ce59187 FEXCore/Interface/Core/JIT: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 1032224d62 FEXCore/Interface/Core/Interpreter: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 44767901fe FEXCore/Interface/Core/Dispatcher: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 1220c86573 FEXCore/Interface/Core/ArchHelpers: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 6ace406a2f FEXCore/Interface/Core/ObjectCache: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 38f1536255 FEXCore/Interface/Core/OpcodeDispatcher: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 324473651e FEXCore/Interface/Core/X86Tables: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 573148b27a FEXCore/Interface/Core/VSyscall: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 12e1c2eaa0 FEXCore/Interface/Context: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek c678ea3060 FEXCore/Interface/GDBJIT: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek e570b07ba0 FEXCore/Interface/Thunks: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek afda4d6b7a FEXCore/Interface/IR: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 22daa506f6 FEXCore/Interface/Config: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek e85b90c614 FEXCore/Common: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 9d3d33fa27 FEXCore/Utils: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 0d9dce987d Merge pull request #3126 from neobrain/feature_better_wayland_thunks64
Thunks/wayland: Add support for APIs required by zink and Super Meat Boy
2023-09-19 10:44:53 -07:00
Ryan Houdek 65d558b2c4 Merge pull request #3119 from alyssarosenzweig/opt/x87-sel
Make x87 FCMOV slightly less terrible
2023-09-19 10:34:29 -07:00
Tony Wasserka b00d413961 Thunks/wayland: Add more message signatures required by Super Meat Boy with zink 2023-09-19 17:33:24 +02:00
Tony Wasserka 6b54540756 Thunks/wayland: Add support for message signatures with nullable arguments 2023-09-19 17:33:24 +02:00
Tony Wasserka 356a42d330 Thunks/wayland: Reorder listener signatures alphabetically 2023-09-19 17:33:24 +02:00
Tony Wasserka 8fcf419183 Thunks/wayland: Add more functions required by Super Meat Boy via libdecor and SDL 2023-09-19 17:33:23 +02:00
Alyssa Rosenzweig 83c8b64c50 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig 25943d1d17 OpcodeDispatcher: Sigh.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig bf03dab295 Arm64: Use csetm
Saves some moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig 8adfaa9aa6 OpcodeDispatcher: Use SelectCC for x87
Better code gen and will benefit from future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:37:54 -04:00
Ryan Houdek ca6570d5de FEXCore/Include: Adds SPDX identifier 2023-09-18 22:13:10 -07:00
Ryan Houdek 3026f7249c FEX/Tools/CommonTools/Linux: Adds SPDX identifier 2023-09-18 22:03:29 -07:00
Ryan Houdek bea29fd2ba FEX: Moves some Linux utils to CommonTools
Was kind of in a weird place before.
2023-09-18 22:01:56 -07:00
Ryan Houdek fc55091fc5 FEX/Common: Adds SPDX identifier 2023-09-18 21:52:20 -07:00
Ryan Houdek 782cf3f7c7 Tools/Opt: Remove. Unused. 2023-09-18 21:45:25 -07:00
Ryan Houdek 8f25e9d3e6 FEXLoader: Adds SPDX identifier 2023-09-18 21:44:23 -07:00
Ryan Houdek 01175e2e7c FEXLoader/LinuxSyscalls: Adds SPDX identifier 2023-09-18 21:43:17 -07:00
Ryan Houdek 5d9d539495 FEXLoader/LinuxSyscalls/Utils: Adds SPDX identifier 2023-09-18 21:38:13 -07:00
Ryan Houdek efb5624db6 FEXLoader/LinuxSyscalls/EmulatedFiles: Adds SPDX identifier 2023-09-18 21:37:51 -07:00
Ryan Houdek fe0a16f478 FEXLoader/HostRunner: Adds SPDX identifier 2023-09-18 21:37:02 -07:00
Ryan Houdek d9d376d40d FEXLoader/ArchHelpers: Adds SPDX identifier 2023-09-18 21:36:23 -07:00
Ryan Houdek 74e7f88449 FEXLoader/AOT: Adds SPDX identifier 2023-09-18 21:35:59 -07:00
Ryan Houdek b2811ffc59 FEX/CommonGUI: Adds SPDX identifier 2023-09-18 21:35:25 -07:00
Ryan Houdek f08e1da577 FEX/CommonTools: Adds SPDX identifier 2023-09-18 21:35:07 -07:00
Ryan Houdek e863eba364 FEXGetConfig: Adds SPDX identifier 2023-09-18 21:34:42 -07:00
Ryan Houdek 10081595af FEXGDBReader: Adds SPDX identifier 2023-09-18 21:34:22 -07:00
Ryan Houdek 75d53725e5 FEXConfig: Adds SPDX identifier 2023-09-18 21:33:02 -07:00
Ryan Houdek d21335be85 FEXBash: Adds SPDX identifier 2023-09-18 21:32:25 -07:00
Ryan Houdek e0385cd807 FEXRootFSFetcher: Adds SPDX identifier 2023-09-18 21:31:51 -07:00
Ryan Houdek e962462e79 FEXServer: Adds SPDX identifier 2023-09-18 21:31:18 -07:00
Ryan Houdek 5896c30954 CodeSizeValidation: Adds SPDX identifier 2023-09-18 21:30:30 -07:00
Ryan Houdek 745729cdc2 SoftFloat-3e: Adds preserve_all attribute to all functions used
This will let FEX's JIT be more optimal
2023-09-18 17:42:48 -07:00
Ryan Houdek 95e5d37e4c FEXCore: Adds compile time check support for preserve_all 2023-09-18 17:09:54 -07:00
Ryan Houdek 838293c2f0 FEXCore: Remove unused FallbackhandlerIndex LoadFCW
We removed this once passing in FCW explicitly.
2023-09-18 17:06:46 -07:00
Ryan Houdek da21fc937b FHU: Fixes syscall helper caching
check_cxx_source_compiles caches by variable name, so `compiles` was
getting cached and breaking future checks.
2023-09-18 17:05:40 -07:00
Alyssa Rosenzweig 3b188b7f49 Merge pull request #3118 from Sonicadvance1/spdx_fhu
FHU: Prepend SPDX identifier
2023-09-18 18:48:30 -04:00
Ryan Houdek 94bbd415a2 FHU: Prepend SPDX identifier
Added with `sed -i '1 i\\/\/ SPDX-License-Identifier: MIT' *.h`
2023-09-18 11:45:18 -07:00
Ryan Houdek 2ea2300408 Merge pull request #3110 from Sonicadvance1/buffered_jit_symbols
FEXCore/JitSymbols: Buffer writes to reduce overhead
2023-09-18 11:38:06 -07:00
Ryan Houdek 000fb2efae Merge pull request #3068 from neobrain/feature_thunk_testlib
unittests: Add test thunk library
2023-09-18 10:31:42 -07:00
Ryan Houdek 8b523082af Merge pull request #3116 from alyssarosenzweig/minor/flag-opts
Minor/flag opts
2023-09-18 10:28:51 -07:00
Alyssa Rosenzweig 5d2a3cd322 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 11:01:46 -04:00
Alyssa Rosenzweig df3833edbe OpcodeDispatcher: Use plain Lshl for flags
If we have PF but no CF this simplifies the IR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 11:01:46 -04:00
Tony Wasserka 527b65648f unittests: Enable logging to stderr when invoking FEXLoader 2023-09-18 16:53:35 +02:00
Tony Wasserka bef64c53f8 unittests: Add test thunk library 2023-09-18 16:53:35 +02:00
Alyssa Rosenzweig 8edcd31404 OpcodeDispatcher: Avoid inverting PF
..if we can fold the invert into the reader.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 10:35:39 -04:00
Ryan Houdek fd1b639ad9 Merge pull request #3115 from lioncash/sqxtun
Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
2023-09-17 14:51:27 -07:00
Ryan Houdek 950a8dbfe7 Merge pull request #3114 from lioncash/ins
Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
2023-09-17 14:40:38 -07:00
Lioncache 26e4d8ad59 Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
If the destination and lower data alias, we can
avoid needing to move into a temporary.
2023-09-17 17:37:36 -04:00
Lioncache d54f590b14 Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
If none of the vectors alias the destination, then we can eliminate
an extra move and usage of a temporary.
2023-09-17 16:56:23 -04:00
Ryan Houdek b3269f20ef Merge pull request #3113 from lioncash/shrn
Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
2023-09-17 13:03:42 -07:00
Lioncache 047646be6d Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
In the event the destination and lower source are the same, then
we don't need to perform any moves.
2023-09-17 15:49:54 -04:00
Ryan Houdek 8168a49d10 Merge pull request #3112 from lioncash/assert
Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
2023-09-17 12:40:47 -07:00
Lioncache 7f2fd4e9a0 Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
Ensures that if our temp vectors change in the future that this is
caught at compile-time rather than runtime.
2023-09-17 15:22:39 -04:00
Lioncache 4ea9f08425 ARMEmitter: Mark index and conversion ops as constexpr
Will be used for assertions. Also makes registers more flexible
for compile-time stuff in general.
2023-09-17 15:19:44 -04:00
Ryan Houdek ffb58761c1 Merge pull request #3111 from lioncash/shift
Arm64/VectorOps: Fix SVE aliasing-path  move in VSShr
2023-09-17 12:15:49 -07:00
Lioncache 8ecdb341e2 Arm64/VectorOps: Fix SVE aliasing-path move in VSShr
Seems like this was a typo from 8d11073, since we'd be moving
into a temporary and then never use it.
2023-09-17 14:43:05 -04:00
Ryan Houdek 0c5c146fcf FEXCore/JitSymbols: Buffer writes to reduce overhead
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.

The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.

To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
  clang benchmark and seeing how frequently it was still writing.
   - 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
  symbols.
   - In most cases 100ms is enough that you won't notice the blip in
     perf.

One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
2023-09-16 17:52:46 -07:00
Ryan Houdek ad8b0c673f Merge pull request #3109 from lioncash/shlx
OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
2023-09-15 18:49:36 -07:00
Ryan Houdek e574cfe681 Merge pull request #3108 from lioncash/mulx
OpcodeDispatcher: Improve output of MULX
2023-09-15 18:09:02 -07:00
Lioncache e9be291cec OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
We can remove some unnecessary moves for the 32-bit cases and
collapse the operations down to a single instruction.
2023-09-15 21:05:50 -04:00
Lioncache d4f87c7db1 OpcodeDispatcher: Improve output of MULX
We can cut down on a few of the generated moves. For
the case where both destinations alias one another,
we can just calculate the high part instead of both of them.
2023-09-15 20:52:02 -04:00
Ryan Houdek 4604c01986 Merge pull request #3107 from lioncash/pext
Arm64/ALUOps: Remove spills in PEXT
2023-09-15 17:40:44 -07:00
Lioncache b0c8ff0ea6 Arm64/ALUOps: Remove spills in PEXT
Reduces the number of emitted instructions for a
corresponding PEXT instruction.

We no longer spill for this IR op.
2023-09-15 19:39:51 -04:00
Ryan Houdek 647629ac23 Merge pull request #3105 from lioncash/rorx
OpcodeDispatcher: Handle RORX corner cases better
2023-09-15 14:55:17 -07:00
Lioncache be90e76422 Arm64/ALUOps: mov in the case of full 32-bit/64-bit BFE
Allows register-renaming mechanisms to be invoked more frequently
2023-09-15 17:38:01 -04:00
Lioncache 4a37ea4819 OpcodeDispatcher: Handle RORX corner cases better
There are a few cases where we were emitting code when we
didn't really need to, or could emit less.
2023-09-15 17:36:36 -04:00
Ryan Houdek 6e08ac65b9 Merge pull request #3106 from lioncash/clwb
HostFeatures: Fix x86 CLWB support check
2023-09-15 14:06:54 -07:00
Lioncache 8705de1893 HostFeatures: Fix x86 CLWB support check
This was clobbering the BMI2 boolean unintentionally.
2023-09-15 16:38:45 -04:00
Alyssa Rosenzweig c8e7c347c3 Merge pull request #3100 from Sonicadvance1/optimize_cmov
OpcodeDispatcher: Optimize cmov
2023-09-15 15:17:51 -04:00
Ryan Houdek 3d0b66407e Merge pull request #3104 from lioncash/vperm2
InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
2023-09-15 11:27:11 -07:00
Lioncache c86b6dc690 InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
Allows viewing the codegen for cases where conditional zeroing is performed.

Also fixes up the vperm2f variants shorthanding one of the registers
to make everything a little more explicit.
2023-09-15 14:10:11 -04:00
Ryan Houdek 773e9465bc Merge pull request #3103 from lioncash/warn
DeadContextStoreElimination: Silence unused function warning
2023-09-15 10:57:40 -07:00
Lioncache e1ed7f43fd DeadContextStoreElimination: Turn LastAccessType into an enum class
Makes the type stricter in terms of implicit conversions.
2023-09-15 13:23:00 -04:00
Ryan Houdek 3eb501aa27 InstCountCI: Update for optimized NZCV and cmov 2023-09-15 10:11:50 -07:00
Ryan Houdek d5b58eebaf OpcodeDispatcher: Optimize cmov
cmov was quite terrible in its implementation. Some things of note:
- NZCV cache would cause store for no reason
- {16,32}-bit would zero extend sources for no reason
- 16-bit would zero extend result for no reason

A bunch of flag testing is still doing a ubfx plus compare against zero
when it could end up being a tst instead, but this is a step in the
right direction and switches over to explicit sized selects.
2023-09-15 10:09:37 -07:00
Ryan Houdek 6dbbd9ecfc OpcodeDispatcher: Duplicate SelectCC but with Explicit result size
This is a temporary measure as we are moving Select operations over to
explicit sizes. Once we remove all uses of SelectCC then it will get
removed.
2023-09-15 10:09:37 -07:00
Ryan Houdek 759cc0025a OpcodeDispatcher: Add a dirty flag for tracking NZCV status
Cached NZCV reads don't need to be written back at the end of the block.
This will remove one instruction from the end of some blocks.
2023-09-15 10:09:37 -07:00
Lioncache d05f890147 DeadContextStoreElimination: Silence unused function warning 2023-09-15 13:08:32 -04:00
Alyssa Rosenzweig 9152fb030e Merge pull request #3102 from alyssarosenzweig/inline-xor
Inline constant with PF calculation
2023-09-15 12:48:33 -04:00
Alyssa Rosenzweig 2385c275ac InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:33:53 -04:00
Alyssa Rosenzweig d29b8bab36 OpcodeDispatcher: Inline constant in PF calculation
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:33:53 -04:00
Ryan Houdek b6922dff57 Merge pull request #3101 from alyssarosenzweig/opt/dec
Optimize out carry invert for DEC
2023-09-15 09:30:08 -07:00
Alyssa Rosenzweig c560a88de4 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:10:57 -04:00
Alyssa Rosenzweig fc02f38435 IR: Only invert CF for NZCV if needed
If we are going to throw away the updated value of CF anyway there is no point
wasting an instruction to invert CF. Add an IR toggle for that so the arm64 JIT
can make better choices.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:08:22 -04:00
Ryan Houdek d5782567e8 Merge pull request #3077 from Sonicadvance1/x86_shifted
FEXCore: Implements support for shifted bitwise ops
2023-09-15 08:09:35 -07:00
Ryan Houdek 060433621a Merge pull request #3097 from Sonicadvance1/disable_enhanced_tso
FEXCore: Disable Enhanced REP MOVSB if Atomic TSO is enabled
2023-09-15 08:08:36 -07:00
Ryan Houdek 9866e238d5 Merge pull request #3080 from Sonicadvance1/defer_softfloat
FEXCore: Defer setting x87 softflow rounding mode until use
2023-09-15 08:08:04 -07:00
Mai f5c4e28696 Merge pull request #3098 from Sonicadvance1/optimize_vectors_sve
Arm64: Optimize wide shifts slightly for 64-bit OpSize
2023-09-15 05:21:20 -04:00
Mai 96bbd01ad6 Merge pull request #3096 from Sonicadvance1/optimal_crc
OpcodeDispatcher: Optimize CRC32
2023-09-15 05:19:05 -04:00
Mai f84a264b0e Merge pull request #3095 from Sonicadvance1/bswap
OpcodeDispatcher: Optimize 16-bit MOVBE
2023-09-15 05:18:26 -04:00
Mai a8c17201b5 Merge pull request #3099 from Sonicadvance1/explicit_but_implicit_select
IR: Changes Select operation to not have implicit sizes
2023-09-15 05:17:47 -04:00
Ryan Houdek d81d89c4fb IR: Changes Select operation to not have implicit sizes
Changes the helper which all the source uses to still calculate the size
implicitly. This is going to take a while to convert all implicit uses
over to the explicit operation.

Get us started by at least having the IR operation itself be explicit.
2023-09-14 20:48:16 -07:00
Ryan Houdek 92212c48f1 IR: Fixes parsing of default arguments with colons
We need to split on the first colon, not every colon in the arguments.

This will be used in the next changes.
2023-09-14 20:37:45 -07:00
Ryan Houdek 42a24bbbd1 InstCountCI: Update CI for optimized wide shifts 2023-09-14 19:53:18 -07:00
Ryan Houdek 021c99e233 Arm64: Optimize wide shifts slightly for 64-bit OpSize
Wide shifts under SVE use 64-bit source elements. If a smaller element
overlaps the 64-bit shift element then it uses that shift
eg:
- Src1[15:0] >> Shift[63:0]
- Src1[31:16] >> Shift[63:0]
- Src1[47:32] >> Shift[63:0]
- Src1[63:48] >> Shift[63:0]
- After this point it will switch to the next 64-bit shift element
- Src1[79:64] >> Shift[127:64]
- Src1[95:80] >> Shift[127:64]
- Src1[111:96] >> Shift[127:64]
- Src1[127:112] >> Shift[127:64]

As seen, we can skip the duplication of the scalar element if the OpSize
is 64-bit, this makes MMX emulation slightly more optimal here.
This also means that a few instructions that weren't claimed to be
optimal actually are since they need the duplication operation (which
vixl always labels as a mov).
2023-09-14 19:35:56 -07:00
Ryan Houdek f730339365 OpcodeDispatcher: Reorder vector loads in shifts
This affects codegen due to RA quirks. This ensures that the wide shifts
don't have to generate a movprfx.
2023-09-14 19:34:00 -07:00
Ryan Houdek 40a4eb90af FEXCore: Disable Enhanced REP MOVSB if TSO is enabled
Hades and the vcruntime hits this very hard in memmove.

`86.56%  [JIT] tid 458574        [.] JIT_0x18000c375_0x7fffc94790c8`

```asm
   0x00007fffc94790f8:  ldaprb  w3, [x2]
   0x00007fffc94790fc:  stlrb   w3, [x1]
   0x00007fffc9479100:  add     x1, x1, #0x1
   0x00007fffc9479104:  add     x2, x2, #0x1
   0x00007fffc9479108:  sub     x0, x0, #0x1
   0x00007fffc947910c:  cbnz    x0, 0x7fffc94790f8
```

This performance is terrible because Cortex's LRCPC performance is bottom-tier.
Work around the performance issue by forcing things to do larger moves with vector moves instead.
2023-09-14 17:23:02 -07:00
Ryan Houdek 31ad26202e InstCountCI: Update for Optimized CRC32 2023-09-14 16:23:33 -07:00
Ryan Houdek a3115d4699 OpcodeDispatcher: Optimize CRC32
The only version of this instruction that was generating optimal code
was the one with 64-bit destination and source.

Optimizes the rest of the operating sizes so that they are all optimal
at one instruction translations
2023-09-14 16:21:42 -07:00
Ryan Houdek c3ead80927 InstCountCI: Update for optimized 16-bit movbe 2023-09-14 16:05:38 -07:00
Ryan Houdek 80cda1bb18 OpcodeDispatcher: Optimize 16-bit MOVBE
16-bit MOVBE is a bit of a special case where it loads 16-bits in to the
bottom of the GPR without clearing the upper bits of the register.
Which means 32-bits or 64-bits depending on operating mode.

Arm64 doesn't support a 16-bit bswap so it needs to operate at 32-bits
instead. We then can insert the resulting bits of the 32-bit rev with a
bfxil in to the lower bits of the resulting destination register.

This allows 16-bit movbe to be optimal now.
2023-09-14 16:05:16 -07:00
Ryan Houdek 90ddee5f8d IR: Implements support for arm64 bfxil
This is useful for extracting a width from a register and inserting in
to the lower bits of a destination.
2023-09-14 15:52:55 -07:00
Ryan Houdek 6fdf2f963b Merge pull request #3082 from Sonicadvance1/minor_storeregsra_opt
FEXCore: Minor optimization to StoreRegisterSRA
2023-09-14 14:54:39 -07:00
Ryan Houdek e1eb151051 Merge pull request #3094 from neobrain/refactor_reorder_ci
CI: Run tests with <30s runtime first
2023-09-14 13:56:12 -07:00
Tony Wasserka 3f8bf01f75 CI: Run tests with <30s runtime first 2023-09-14 20:46:50 +02:00
Mai 92824f5e4d Merge pull request #3093 from Sonicadvance1/optimize_blendp
OpcodeDispatcher: Optimize blendp{s,d}
2023-09-14 00:33:00 -04:00
Mai 213d3c4e2b Merge pull request #3091 from Sonicadvance1/optimize_pinsr
OpcodeDispatcher: Optimize pins{b,w,d,q}
2023-09-14 00:30:56 -04:00
Mai d4c6749d2a Merge pull request #3090 from Sonicadvance1/optimize_pextr
OpcodeDispatcher: Optimize pextr{b,w}
2023-09-14 00:30:43 -04:00
Mai 1804b007ec Merge pull request #3092 from Sonicadvance1/instcountci_compile_log
InstCountCI: Add log before compiling instruction
2023-09-13 23:14:42 -04:00
Mai 655cee070d Merge pull request #3089 from Sonicadvance1/optimize_pshufd
OpcodeDispatcher: Optimize shufpd
2023-09-13 23:12:10 -04:00
Ryan Houdek 28309a1cc5 InstCountCI: Update for blend 2023-09-13 20:08:03 -07:00
Ryan Houdek 29f824cf7a OpcodeDispatcher: Optimize blendp{s,d}
Optimal blendps is worst case 2 instructions.
FEX's RA doesn't quite get there since it can't see through multiple
instructions with SRA destinations. That'll be fixed in the future.

Optimal blendps is always one instruction, one is a no-op.
We always hit this.
2023-09-13 20:06:39 -07:00
Ryan Houdek 2bafa2c26f InstCountCI: Update for optimized pins{b,w,d,q} 2023-09-13 19:53:05 -07:00
Ryan Houdek 5e7d793a6a OpcodeDispatcher: Optimize pins{b,w,d,q}
Inserting from a GPR and memory can both be optimized. These are now
optimal

Needs #3088 merged first.
2023-09-13 19:53:05 -07:00
Ryan Houdek 3e40713ccc InstCountCI: Update for optimized pextr{b,w} 2023-09-13 19:51:56 -07:00
Ryan Houdek 33a2fbb896 OpcodeDispatcher: Optimize pextr{b,w}
Cleans up the code which had special cased some 32-bit optimization
which is unnecessary now that both 8-bit and 16-bit are also optimized.

When FEX does a VExtractToGPR, the result is zero extended to the full
GPR register size. This means we don't need to do a zero extend when
storing to a guest GPR.

Makes pextr{b,w} optimal now.

Needs #3088 merged first.
2023-09-13 19:51:56 -07:00
Ryan Houdek 853ded7df7 InstcountCI: Update for optimized shufpd 2023-09-13 19:50:15 -07:00
Ryan Houdek 67914157cb OpcodeDispatcher: Optimize shufpd
This one is very satisfying since there are only four variants and each
one of them converts to a single instruction.

Needs #3088 merged first
2023-09-13 19:50:15 -07:00
Mai 750d90939d Merge pull request #3088 from Sonicadvance1/instcountci_missing_secondary_opsize
InstCountCI: Adds missing instructions from Secondary OpSize tables
2023-09-13 22:49:17 -04:00
Mai 31d828390f Merge pull request #3087 from Sonicadvance1/tbl2_implementation
OpcodeDispatcher: Implement shufps with VTBL2 in worst case
2023-09-13 22:48:53 -04:00
Ryan Houdek 6c2f8ab085 InstCountCI: Add log before compiling instruction
If CI faults out due to a bug then we would have no log as to which
instruction caused the issue.

I find myself adding this each time an assert fires to see what
instruction it was working on. Just add it directly.
2023-09-13 14:33:33 -07:00
Ryan Houdek 2aea401189 InstCountCI: Adds missing instructions from Secondary OpSize tables
I managed to miss a whole section of instructions from the secondary
opsize tables. This resulted in four instructions missing from the
database.

Adds cmppd, pinsrw, pextrw, and shufpd which are all non-optimal
instruction implementations.
2023-09-13 11:48:50 -07:00
Ryan Houdek c008671509 unittests/asm: Add test with inverted sources
To ensure this is tested with non sequential source registers.
2023-09-13 11:31:20 -07:00
Ryan Houdek 5903be156c InstCountCI: Update for shufps tbl opt 2023-09-13 11:31:20 -07:00
Ryan Houdek db5056f275 OpcodeDispatcher: Implement shufps with VTBL2 in worst case
In the case that source registers are sequential then this turns in to a
load of the vector constant (2 instructions) and the single tbl
instruction.

If the registers aren't sequential then the tbl turns in to 2 moves and
then the single tbl, which with zero-cycle rename isn't too bad.

Since this is a worst case option this is significantly better than the
previous implementation doing a bunch of inserts which was always 9
instructions.
We should still strive to implement faster versions without the use of
TBL2 if possible but this makes it less of a concern.
2023-09-13 11:31:20 -07:00
Ryan Houdek e9d96ce538 IR: Implements support for VTBL2
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.

The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
2023-09-13 11:31:20 -07:00
Ryan Houdek 444d4c082d Int: Fixes typo in LoadNamedVectorIndexedConstant
Surprising this didn't break anything before this.
2023-09-13 11:31:20 -07:00
Ryan Houdek cfe620ab15 Merge pull request #3085 from Sonicadvance1/optimize_shufps
OpcodeDispatcher: Optimize a bunch of shufps variants
2023-09-12 21:53:33 -07:00
Ryan Houdek ea8d63350a InstCountCI: Updates for optimized shufps 2023-09-12 19:58:07 -07:00
Ryan Houdek e37cef8283 unittests: Implement shufps optimization test
Tests all current forms of shufps optimizations.
2023-09-12 19:58:07 -07:00
Ryan Houdek 3f1979286f OpcodeDispatcher: Optimize a bunch of shufps variants
Hits a whole bunch of common cases, most of which then emit optimal code
generation.
Two cases that use VInsElement hit the RA quirk where the SRA
destination is dead but RA doesn't see it, so it ends up doing a couple
moves. If RA gets fixed then those two moves will go away.

There are definitely still cases that we could emit more optimal code.
Additionally we could implement a TBL2 IR operation to do a LUT approach
for ones we don't cover.

Problem with implementing a TBL2 ir operation is that we have no way to
ensure registers are sequential so we would need to always do moves
```asm
ldr v2, <LUT Table>
mov v0, v16
mov v1, v18
tbl v16.16b, { v0.16b, v1.16b }, v2.16b
```

Which to be fair isn't terrible, and if we're lucky that the guest uses
sequential registers we can naturally get the more optimal code path.
Ideally our RA could push some operations in to sequential registers but
that's not possible currently.

I'll do a follow-up PR that implements TBL2.
2023-09-12 19:58:07 -07:00
Ryan Houdek d5c3036bc2 JITx86: Fixes VREV64 with 32-bit element size.
This has been incorrect since it has been implemented.
Noticed when implementing optimizations.
2023-09-12 19:23:59 -07:00
Mai ebdca02218 Merge pull request #3084 from Sonicadvance1/optimize_bswap
OpcodeDispatcher: Optimize 32-bit bswap
2023-09-12 20:09:00 -04:00
Mai dda5861bdd Merge pull request #3081 from Sonicadvance1/fix_waitpid
Tools: Fixes usage of waitpid in the face of EINTR
2023-09-12 19:35:05 -04:00
Mai f7e652b616 Merge pull request #3083 from Sonicadvance1/optimize_nop_move
OpcodeDispatcher: Optimize NOP vector move
2023-09-12 19:34:36 -04:00
Ryan Houdek 65bc159ff1 InstCountCI: Update for bswap optimization 2023-09-12 16:19:55 -07:00
Ryan Houdek c362d3a9d8 OpcodeDispatcher: Optimize 32-bit bswap
Removes a redundant move, making it optimal now.
2023-09-12 16:19:10 -07:00
Ryan Houdek 8a44be0c30 InstCountCI: Update for NOP vector moves
Adds a couple of instructions that get tested in this code path.
2023-09-12 16:11:43 -07:00
Ryan Houdek 304dba5f20 OpcodeDispatcher: Optimize NOP vector move
Move instruction to itself here is a nop.
Need to be careful about AVX operations which use a different handler
since those might actually zero the upper bits on 128-bit move
2023-09-12 16:10:39 -07:00
Ryan Houdek b2e61d2deb InstCountCI: Update for minor storeregistersra opt 2023-09-12 14:38:46 -07:00
Ryan Houdek e6c0bebee9 FEXCore: Minor optimization to StoreRegisterSRA
{Load,Store}RegisterSRA always loads or stores GPRSize. 8-bit and 16-bit
are vestigial and all OpcodeDispatcher usage will load the full GPR size
(32-bit or 64-bit) and then extract or insert as necessary.

This cleans up a few bits of codegen in InstCountCI.
2023-09-12 14:36:08 -07:00
Ryan Houdek aa017116b3 Tools: Fixes usage of waitpid in the face of EINTR
waitpid can return early if interrupted due to EINTR.
Loop on this case and try again.
2023-09-12 12:41:43 -07:00
Ryan Houdek 97a6184e53 InstCountCI: Update for FCW optimization 2023-09-12 05:21:06 -07:00
Ryan Houdek 76bd81af15 FEXCore: Defer setting x87 softflow rounding mode until use
Currently FEX will always jump out of the JIT any time FCW was getting
written to, ensuring that the softfloat state is setup to rounding at
the time of FCW getting written.
This has the unintended side-effect that even in "x87 reduced precision"
mode we were jumping out of the JIT.
This hit a real world use case of an installer reloading FCW after every
x87 operation and generating a block with 2297 instructions.

Instead when jumping out of the JIT for handling x87 operations, load
FCW and pass it as the first argument of the handler. Setting the
softfloat state at that point.

This helps the installer's hottest block by cutting it down to 1477
instructions. 64.3% of the original size. The code block is still
burning 90% of the CPU time of the installer but the performance is
significantly better while it is doing its decompression.

In order to optimize this installer's block of code more then we will
likely need to optimize out x87 stack usage.
2023-09-12 05:21:06 -07:00
Mai 90f7937146 Merge pull request #3079 from Sonicadvance1/recover_two_temps
Arm64: Recover two unused vector vector temporary registers
2023-09-11 22:06:03 -04:00
Mai 98f148766d Merge pull request #3078 from Sonicadvance1/detect_flagm
HostFeatures: Detect FlagM/2
2023-09-11 20:57:43 -04:00
Ryan Houdek 9c44e295fa InstCountCI: Update for recovering two vector temps
All the changes are RA changes and spilling/filling taking another
instruction.
2023-09-11 16:50:52 -07:00
Ryan Houdek b5a1d323c2 Arm64: Recover two unused vector vector temporary registers
This leaves us with two temporary vectors that the JIT can use.
As of last month we stopped using v2 and v3 as temporaries and these can
now be given back to the JIT.

Ensures that the registers are still sequentially ordered and adds
support for spilling the FPR counts that are aligned by 2 instead of 4.
Adds a couple of instructions to filling and spilling but isn't that big
of an issue.

InstcountCI has some ridiculously large changes just because RA is
starting at a new register number.
2023-09-11 16:48:25 -07:00
Ryan Houdek b453439968 HostFeatures: Detect FlagM/2
Currently unused but at least detect the feature so that our Arm64 JIT
can use it in the future.
2023-09-11 16:41:30 -07:00
Ryan Houdek 863331b117 FEXCore: Implements support for shifted bitwise ops
This wasn't implemented initially for the interpreter and x86 JIT.

This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.

The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
2023-09-11 13:17:35 -07:00
Mai 48521a4416 Merge pull request #3075 from Sonicadvance1/optimize_bt_ops
OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
2023-09-11 16:05:33 -04:00
Mai 6fe643d270 Merge pull request #3076 from Sonicadvance1/enable_enhanced_rep_movs
CPUID: Enabled Enhanced REP MOVSB/STOSB
2023-09-11 15:35:56 -04:00
Mai fbc4bda7a6 Merge pull request #3074 from Sonicadvance1/hwcap2_fsgsbase
ELFCodeLoader: Expose FSGSBase in getauxval HWCAP2
2023-09-11 15:35:26 -04:00
Mai 6d9b52452e Merge pull request #3072 from Sonicadvance1/crc32_is_fixed_size
IR: Changes crc32 operation to always return a 32-bit result.
2023-09-11 15:34:37 -04:00
Mai 950007c815 Merge pull request #3071 from Sonicadvance1/update_rcl_opsize
OpcodeDispatcher: Update 32/64-bit RCL for operating size
2023-09-11 15:34:06 -04:00
Mai d029394c27 Merge pull request #3070 from Sonicadvance1/update_rcr_opsize
OpcodeDispatcher: Update 32/64-bit RCR for operating size
2023-09-11 15:33:34 -04:00
Mai 879fcdc6fe Merge pull request #3069 from Sonicadvance1/fix_redundant_load_rclse
IR:RCLSE: Partially reenables the RCLSE pass
2023-09-11 15:32:55 -04:00
Ryan Houdek 2f77982b54 CPUID: Enabled Enhanced REP MOVSB/STOSB
Missed with #2490.
This changes behaviour of glibc's memmove slightly, seems to recover a
bit of performance on Half-Life 2's title screen.
2023-09-10 20:49:08 -07:00
Ryan Houdek e3a00fb2fb InstCountCI: Update for BT minor opt 2023-09-10 20:23:08 -07:00
Ryan Houdek 4feb059f51 OpcodeDispatcher: Optimize the case of all flags invalidated
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
2023-09-10 20:16:29 -07:00
Ryan Houdek 3d1bbe505d OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
These instructions set all the flags to undefined and moves the
resulting bit in to CF. No need to calculate the deferred flags when
we are about to write over them.
2023-09-10 20:16:29 -07:00
Ryan Houdek b2a42b6c61 ELFCodeLoader: Expose FSGSBase in getauxval HWCAP2
We have supported this since #163 but we haven't been exposing the
feature in hwcap2.

We have exposed it in CPUID this entire time, just not in hwcap2.
2023-09-10 17:00:21 -07:00
Ryan Houdek 315d1855de IR: Changes crc32 operation to always return a 32-bit result.
CRC32 is always a 32-bit sized operation even with a 64-bit source
value.
This doesn't change any InstCountCI results.
2023-09-09 10:02:30 -07:00
Ryan Houdek 93246878e2 InstCountCI: Update for rcl explicit size change 2023-09-09 09:40:36 -07:00
Ryan Houdek 6c62691af0 OpcodeDispatcher: Update 32/64-bit RCL for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:40:12 -07:00
Ryan Houdek ee5aed51d8 InstCountCI: Update for rcr expliti size change 2023-09-09 09:35:13 -07:00
Ryan Houdek 47f50a7008 OpcodeDispatcher: Update 32/64-bit RCR for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:33:27 -07:00
Ryan Houdek be07254935 Merge pull request #3067 from neobrain/refactor_thunks
Thunks: Minor restructuring and small cleanups
2023-09-07 20:16:17 -07:00
Ryan Houdek 636f8aa4a7 Arm64: Fix undefined behaviour in Push operation
Arm64 store with writeback when source register is the same register as
the address is undefined behaviour.
Depending on hardware details this can do a whole bunch of things.

This situation happens when the x86 code does `push rsp` which is quite
common for applications to do. We would then convert this to a `str x8, [x8, #-8]!`
Which results in undefined behaviour.

Now that redundant loads are optimized this showed up as an issue. Adds
a unit test to ensure we don't hit this again.
2023-09-07 17:38:39 -07:00
Ryan Houdek 22ca46a227 Arm64: Fixes SVE V{S,U}MulH
When the destination overlaps one of the sources we must be careful to
follow a movprfx rule.
```
The destination register must not refer to architectural register state
referenced by any other source operand register of this instruction.
```

We ended up in a situation in the vpmulh{u,}w AVX tests where zm was
overlapping the destination which violated that rule. This also
generated invalid code for this instruction.
```
[INFO] movprfx z6, z4
[INFO] umulh z6.h, p6/m, z6.h, z6.h
```

As seen, we were overwriting one of the sources because the destination
overlapped it. Now instead check if each individual overlap so invalid
code isn't generated.

InstCountCI results aren't affected since this only happens in
situations with multiple instructions.
2023-09-07 16:47:08 -07:00
Ryan Houdek 7b80427de0 OpcodeDispatcher: Remove BLENDV "optimization"
Now that the RCLSE pass finally optimizes redundant loads again this
optimization that lives in the OpcodeDispatcher can be removed.

With InstCountCI reran, the pblendvb results don't change at all, as
expected.
2023-09-07 16:00:56 -07:00
Ryan Houdek b753b9ffa2 InstCountCI: Updates for RCLSE fix
Adds two `packsswb` tests to ensure redundant sources are getting
optimized as expected.
2023-09-07 15:58:43 -07:00
Ryan Houdek c62b5a3103 IR:RCLSE: Partially reenables the RCLSE pass
This is taking steps to start fixing RCLSE which was started by #2700.
Same situation as that PR, since #2170 when we converted
{Load,Store}Context in to {Load,Store}Register we broke this pass
entirely. It hasn't been doing anything for redundant GPRs and FPRs
since at least November of last year.

Technically it was potentially still optimizing redundant MMX
accesses, but it is so broken that it doesn't matter.

Instead of going all in like #2700 did, tear down the pass and start
again. We are now /only/ optimizing redundant context/register loads.
This fixes an issue that comes up commonly where the same register used
as sources was getting loaded twice, causing redundant moves.

`packsswb xmm0, xmm0` for example was generating a four instruction
sequence instead of three instructions because we weren't eliminating
the redundant load.

Going to take reimplementing all the optimizations that this pass does
in steps. This way we can track any regression in the independent steps
unlike what happened in #2700.

Confirmed that Proton/Sonic Mania still works after this.
2023-09-07 15:50:27 -07:00
Ryan Houdek 3c729bcacb IR/RCLSE: Removes unused CalculateControlFlowInfo
This is unused and this only optimizes inside of a block.
2023-09-07 15:49:29 -07:00
Tony Wasserka 677b77f1bb Thunks: Simplify PackedArguments invocation code 2023-09-07 13:56:53 +02:00
Tony Wasserka 024fb268c0 unittests/ThunkLibs: Move utility code to a dedicated header 2023-09-07 13:56:53 +02:00
Tony Wasserka ab8bc052a0 Thunks/gen: Split interface parsing and code emission into separate files 2023-09-07 13:56:53 +02:00
Tony Wasserka 971821460c Thunks/gen: Move diagnostic marshaling helper to a dedicated header 2023-09-07 13:56:53 +02:00
Ryan Houdek 615ab8d80c Merge pull request #3066 from neobrain/fix_procfsint_regression
FileManagement: Fix inverted boolean check for procfs/interpreter support
2023-09-06 17:37:12 -07:00
Tony Wasserka 83c74e86c8 FileManagement: Fix inverted boolean check for procfs/interpreter support 2023-09-06 17:00:58 +02:00
Mai 9ff5544d55 Merge pull request #3065 from Sonicadvance1/fix_docs_location
Scripts: Update generate_doc_outline for moved FEXCore
2023-09-06 10:54:20 -04:00
Ryan Houdek 2b9265d9fe Scripts: Update generate_doc_outline for moved FEXCore
Otherwise all of the FEXCore docs get deleted.
2023-09-05 23:23:27 -07:00
634 changed files with 51835 additions and 41762 deletions

No files matched your search

+50 -50
View File
@@ -17,11 +17,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
build_plus_test:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64], [self-hosted, ARMv8.0], [self-hosted, ARMv8.2], [self-hosted, ARMv8.4]]
arch: [[self-hosted, ARMv8.0], [self-hosted, ARMv8.2], [self-hosted, ARMv8.4]]
fail-fast: false
steps:
@@ -65,7 +65,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -78,18 +78,6 @@ jobs:
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -102,30 +90,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gvisor tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gvisor_tests
- name: GVisor Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GVisor.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -150,17 +114,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC32.log || true
- name: Struct verifier tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target struct_verifier
- name: Struct verifier Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -244,6 +197,53 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkResults.log || true
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gvisor tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gvisor_tests
- name: GVisor Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GVisor.log || true
- name: Struct verifier tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target struct_verifier
- name: Struct verifier Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+27 -28
View File
@@ -24,12 +24,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
glibc_fault_test:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
# Run on an x86 device and any ARM runner.
arch: [[self-hosted, x64], [self-hosted, ARM64]]
arch: [[self-hosted, ARM64]]
fail-fast: false
steps:
@@ -73,7 +72,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -86,18 +85,6 @@ jobs:
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -110,18 +97,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -179,6 +154,30 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+107
View File
@@ -0,0 +1,107 @@
name: Hostrunner tests
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
# Customize the CMake build type here (Release, Debug, RelWithDebInfo, etc.)
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
hostrunner_tests:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+31 -3
View File
@@ -16,11 +16,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
instcountci_tests:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64]]
arch: [[self-hosted, x64], [self-hosted, ARM64]]
fail-fast: false
steps:
@@ -64,7 +64,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=False -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=False -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -86,6 +86,25 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_InstCountCI.log || true
- name: Update local repo instcount
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target instcountci_update_tests
- name: Get instcountCI diff
if: ${{ always() }}
shell: bash
working-directory: ${{github.workspace}}/
run: git diff --output=${{runner.workspace}}/build/InstCountCI.diff
- name: Check if InstCountCI Diff exists
if: ${{ always() }}
shell: bash
working-directory: ${{github.workspace}}/
# Check if the file is empty
run: sh -c "! test -s ${{runner.workspace}}/build/InstCountCI.diff"
- name: Truncate test results
if: ${{ always() }}
shell: bash
@@ -107,3 +126,12 @@ jobs:
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
- name: Upload results InstCountCI
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}-instcountci
path: ${{runner.workspace}}/build/InstCountCI.diff
retention-days: 3
+3 -3
View File
@@ -13,11 +13,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
mingw_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64, mingw], [self-hosted, ARM64, mingw]]
arch: [[self-hosted, ARM64, mingw]]
fail-fast: false
steps:
@@ -74,7 +74,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=False -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+1 -1
View File
@@ -16,7 +16,7 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
vixl_simulator:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
+5
View File
@@ -0,0 +1,5 @@
{
"ThunksDB": {
"fex_thunk_test": 1
}
}
+2 -16
View File
@@ -25,7 +25,6 @@ option(ENABLE_JEMALLOC_GLIBC_ALLOC "Enables jemalloc glibc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
option(ENABLE_VIXL_SIMULATOR "Forces the FEX JIT to use the VIXL simulator" FALSE)
@@ -97,11 +96,6 @@ if (ENABLE_GDB_SYMBOLS)
endif()
if (ENABLE_INTERPRETER)
message(STATUS "Interpreter enabled")
add_definitions(-DINTERPRETER_ENABLED=1)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -118,14 +112,6 @@ else()
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
if (NOT ENABLE_X86_HOST_DEBUG)
message(FATAL_ERROR
" Be warned: FEX isn't optimized for x86_64 hosts!\n"
" Support for x86_64 hosts is only for debugging and convenience!\n"
" Don't expect amazing performance or optimal code generation!\n"
" Pass -DENABLE_X86_HOST_DEBUG=True to bypass this message!")
endif()
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
@@ -467,10 +453,10 @@ if (BUILD_THUNKS)
CMAKE_ARGS
"-DBITNESS=64"
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DBUILD_FEX_LINUX_TESTS=${BUILD_FEX_LINUX_TESTS}"
"-DENABLE_CLANG_THUNKS=${ENABLE_CLANG_THUNKS}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_64_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
@@ -485,10 +471,10 @@ if (BUILD_THUNKS)
CMAKE_ARGS
"-DBITNESS=32"
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DBUILD_FEX_LINUX_TESTS=${BUILD_FEX_LINUX_TESTS}"
"-DENABLE_CLANG_THUNKS=${ENABLE_CLANG_THUNKS}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_32_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"Env": "STEAM_GAME_LAUNCH_SHELL=@CMAKE_INSTALL_PREFIX@/bin/FEXBash"
}
}
+6
View File
@@ -144,6 +144,12 @@
"@PREFIX_LIB@/libasound.so.2.0.0"
]
},
"fex_thunk_test": {
"Library": "libfex_thunk_test-guest.so",
"Overlay": [
"@PREFIX_LIB@/libfex_thunk_test.so"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
+16 -9
View File
@@ -13,15 +13,6 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
set(_M_ARM_64 1)
endif()
if (ENABLE_VIXL_SIMULATOR)
# If the vixl simulator is enabled then we are using the ARM64 JIT
option(ENABLE_JIT_X86_64 "Enable the x86_64 JIT" FALSE)
option(ENABLE_JIT_ARM64 "Enable the ARM64 JIT" TRUE)
else()
option(ENABLE_JIT_X86_64 "Enable the x86_64 JIT" ${_M_X86_64})
option(ENABLE_JIT_ARM64 "Enable the ARM64 JIT" ${_M_ARM_64})
endif()
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
set(CMAKE_POSITION_INDEPENDENT_CODE ON)
@@ -33,6 +24,22 @@ set(CMAKE_INCLUDE_CURRENT_DIR ON)
include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
include(CheckCXXSourceCompiles)
set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
check_cxx_source_compiles(
"
__attribute__((preserve_all))
void Testy() {
}
int main() {
return 0;
}"
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
if (HAS_CLANG_PRESERVE_ALL)
message(STATUS "Has clang::preserve_all")
endif ()
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
+1 -1
View File
@@ -144,7 +144,7 @@ def parse_ops(ops):
Argument = Argument.strip()
OpArg = OpArgument()
Split = Argument.split(":")
Split = Argument.split(":", 1)
if len(Split) != 2:
ExitError("Error parsing argument. Missing Type and name colon split")
+20 -54
View File
@@ -107,9 +107,19 @@ set (SRCS
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp
Interface/Core/JIT/Arm64/Arm64Relocations.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -141,7 +151,7 @@ set (SRCS
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/SyscallOptimization.cpp
Interface/IR/Passes/InlineCallOptimization.cpp
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
@@ -159,24 +169,7 @@ if (ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
Utils/AllocatorOverride.cpp)
endif()
if (ENABLE_INTERPRETER)
list(APPEND SRCS
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp)
endif()
set(DEFINES -DTHREAD_LOCAL=_Thread_local)
set(DEFINES -DTHREAD_LOCAL=_Thread_local -DJIT_ARM64)
if (_M_X86_64)
list(APPEND DEFINES -D_M_X86_64=1)
@@ -195,41 +188,14 @@ if (ENABLE_VIXL_DISASSEMBLER)
list(APPEND DEFINES -DVIXL_DISASSEMBLER=1)
endif()
if (ENABLE_JIT_X86_64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
Interface/Core/JIT/x86_64/ALUOps.cpp
Interface/Core/JIT/x86_64/AtomicOps.cpp
Interface/Core/JIT/x86_64/BranchOps.cpp
Interface/Core/JIT/x86_64/ConversionOps.cpp
Interface/Core/JIT/x86_64/EncryptionOps.cpp
Interface/Core/JIT/x86_64/FlagOps.cpp
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp
Interface/Core/JIT/x86_64/x64Relocations.cpp
)
list(APPEND DEFINES -DJIT_X86_64)
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=__attribute__((preserve_all));-DFEXCORE_HAS_PRESERVE_ALL_ATTR=1")
else()
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=;-DFEXCORE_HAS_PRESERVE_ALL_ATTR=0")
endif()
if (ENABLE_JIT_ARM64)
list(APPEND DEFINES -DJIT_ARM64)
list(APPEND SRCS
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp
Interface/Core/JIT/Arm64/Arm64Relocations.cpp
)
endif()
# Some defines for the softfloat library
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ")
set (LIBS fmt::fmt vixl xxhash FEXHeaderUtils)
+1
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/Allocator.h>
+86 -36
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/fextl/fmt.h>
#include "Common/JitSymbols.h"
@@ -26,42 +27,6 @@ namespace FEXCore {
fd = open(PerfMap.c_str(), O_CREAT | O_TRUNC | O_WRONLY | O_APPEND, 0644);
}
void JITSymbols::Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
@@ -86,4 +51,89 @@ namespace FEXCore {
}
}
// Buffered JIT symbols.
void JITSymbols::Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, GuestAddr, CodeSize);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, CodeSize, Name, Offset);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::RegisterNamedRegion(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}\n", HostAddr, CodeSize, Name);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
RegisterNamedRegion(Buffer, HostAddr, CodeSize, Name);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::WriteBuffer(Core::JITSymbolBuffer *Buffer, bool ForceWrite) {
auto Now = std::chrono::steady_clock::now();
if (!ForceWrite) {
if (((Buffer->LastWrite - Now) < Buffer->MAXIMUM_THRESHOLD) &&
Buffer->Offset < Buffer->NEEDS_WRITE_DISTANCE) {
// Still buffering, no need to write.
return;
}
}
Buffer->LastWrite = Now;
auto Result = write(fd, Buffer->Buffer, Buffer->Offset);
if (Result == -1 && errno == EBADF) {
fd = -1;
}
Buffer->Offset = 0;
}
} // namespace FEXCore
+15 -3
View File
@@ -1,5 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <chrono>
#include <cstdint>
#include <cstdio>
#include <memory>
@@ -12,13 +17,20 @@ public:
~JITSymbols();
void InitFile();
void Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
// Allocate JIT buffer.
static fextl::unique_ptr<Core::JITSymbolBuffer> AllocateBuffer() {
return fextl::make_unique<Core::JITSymbolBuffer>();
}
void Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name);
private:
int fd{-1};
void WriteBuffer(Core::JITSymbolBuffer *Buffer, bool ForceWrite = false);
};
}
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
extF80_roundToInt( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
extF80_to_i32( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
extF80_to_i64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t
extF80_to_ui64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t a )
{
union ui128_f128 uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t a )
{
union ui32_f32 uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t a )
{
union ui64_f64 uA;
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t i32_to_extF80( int32_t a )
{
uint_fast16_t uiZ64;
@@ -68,9 +68,11 @@ uint_fast64_t
uint_fast64_t softfloat_roundMToUI64( bool, uint32_t *, uint_fast8_t, bool );
#endif
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t softfloat_roundToI32( bool, uint_fast64_t, uint_fast8_t, bool );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
@@ -109,8 +111,10 @@ float16_t
#define isNaNF32UI( a ) (((~(a) & 0x7F800000) == 0) && ((a) & 0x007FFFFF))
struct exp16_sig32 { int_fast16_t exp; uint_fast32_t sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig32 softfloat_normSubnormalF32Sig( uint_fast32_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t softfloat_roundPackToF32( bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_normRoundPackToF32( bool, int_fast16_t, uint_fast32_t );
@@ -130,8 +134,10 @@ float32_t
#define isNaNF64UI( a ) (((~(a) & UINT64_C( 0x7FF0000000000000 )) == 0) && ((a) & UINT64_C( 0x000FFFFFFFFFFFFF )))
struct exp16_sig64 { int_fast16_t exp; uint_fast64_t sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig64 softfloat_normSubnormalF64Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t softfloat_roundPackToF64( bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_normRoundPackToF64( bool, int_fast16_t, uint_fast64_t );
@@ -155,11 +161,14 @@ float64_t
*----------------------------------------------------------------------------*/
struct exp32_sig64 { int_fast32_t exp; uint64_t sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig64 softfloat_normSubnormalExtF80Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
@@ -181,6 +190,7 @@ extFloat80_t
#define isNaNF128UI( a64, a0 ) (((~(a64) & UINT64_C( 0x7FFF000000000000 )) == 0) && (a0 || ((a64) & UINT64_C( 0x0000FFFFFFFFFFFF ))))
struct exp32_sig128 { int_fast32_t exp; struct uint128 sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig128
softfloat_normSubnormalF128Sig( uint_fast64_t, uint_fast64_t );
@@ -53,6 +53,7 @@ INLINE
uint64_t softfloat_shortShiftRightJam64( uint64_t a, uint_fast8_t dist )
{ return a>>dist | ((a & (((uint_fast64_t) 1<<dist) - 1)) != 0); }
#else
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shortShiftRightJam64( uint64_t a, uint_fast8_t dist );
#endif
#endif
@@ -74,6 +75,7 @@ INLINE uint32_t softfloat_shiftRightJam32( uint32_t a, uint_fast16_t dist )
(dist < 31) ? a>>dist | ((uint32_t) (a<<(-dist & 31)) != 0) : (a != 0);
}
#else
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_shiftRightJam32( uint32_t a, uint_fast16_t dist );
#endif
#endif
@@ -95,6 +97,7 @@ INLINE uint64_t softfloat_shiftRightJam64( uint64_t a, uint_fast32_t dist )
(dist < 63) ? a>>dist | ((uint64_t) (a<<(-dist & 63)) != 0) : (a != 0);
}
#else
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shiftRightJam64( uint64_t a, uint_fast32_t dist );
#endif
#endif
@@ -148,6 +151,7 @@ INLINE uint_fast8_t softfloat_countLeadingZeros32( uint32_t a )
return count;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros32( uint32_t a );
#endif
#endif
@@ -157,6 +161,7 @@ uint_fast8_t softfloat_countLeadingZeros32( uint32_t a );
| Returns the number of leading 0 bits before the most-significant 1 bit of
| 'a'. If 'a' is zero, 64 is returned.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros64( uint64_t a );
#endif
@@ -178,6 +183,7 @@ extern const uint16_t softfloat_approxRecip_1k1s[16];
#ifdef SOFTFLOAT_FAST_DIV64TO32
#define softfloat_approxRecip32_1( a ) ((uint32_t) (UINT64_C( 0x7FFFFFFFFFFFFFFF ) / (uint32_t) (a)))
#else
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecip32_1( uint32_t a );
#endif
#endif
@@ -204,6 +210,7 @@ extern const uint16_t softfloat_approxRecipSqrt_1k1s[16];
| returned is also always within the range 0.5 to 1; thus, the most-
| significant bit of the result is always set.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecipSqrt32_1( unsigned int oddExpA, uint32_t a );
#endif
@@ -240,6 +247,7 @@ INLINE
bool softfloat_le128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{ return (a64 < b64) || ((a64 == b64) && (a0 <= b0)); }
#else
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_le128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
#endif
@@ -255,6 +263,7 @@ INLINE
bool softfloat_lt128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{ return (a64 < b64) || ((a64 == b64) && (a0 < b0)); }
#else
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_lt128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
#endif
@@ -275,6 +284,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftLeft128( uint64_t a64, uint64_t a0, uint_fast8_t dist );
#endif
@@ -296,6 +306,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftRight128( uint64_t a64, uint64_t a0, uint_fast8_t dist );
#endif
@@ -413,6 +424,7 @@ struct uint64_extra
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint64_extra
softfloat_shiftRightJam64Extra(
uint64_t a, uint64_t extra, uint_fast32_t dist );
@@ -492,6 +504,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_add128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
@@ -528,6 +541,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_sub128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
@@ -562,6 +576,7 @@ INLINE struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b )
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b );
#endif
#endif
@@ -570,6 +585,7 @@ struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b );
/*----------------------------------------------------------------------------
| Returns the 128-bit product of 'a' and 'b'.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64To128( uint64_t a, uint64_t b );
#endif
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_add128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_add128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extern const uint16_t softfloat_approxRecip_1k0s[16];
extern const uint16_t softfloat_approxRecip_1k1s[16];
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecip32_1( uint32_t a )
{
int index;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extern const uint16_t softfloat_approxRecipSqrt_1k0s[];
extern const uint16_t softfloat_approxRecipSqrt_1k1s[];
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecipSqrt32_1( unsigned int oddExpA, uint32_t a )
{
int index;
@@ -44,6 +44,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| floating-point NaN, and returns the bit pattern of this value as an unsigned
| integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr )
{
struct uint128 uiZ;
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| Converts the common NaN pointed to by `aPtr' into a 128-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToF128UI( const struct commonNaN *aPtr )
{
struct uint128 uiZ;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| Converts the common NaN pointed to by `aPtr' into a 32-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast32_t softfloat_commonNaNToF32UI( const struct commonNaN *aPtr )
{
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| Converts the common NaN pointed to by `aPtr' into a 64-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t softfloat_commonNaNToF64UI( const struct commonNaN *aPtr )
{
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#define softfloat_countLeadingZeros32 softfloat_countLeadingZeros32
#include "primitives.h"
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros32( uint32_t a )
{
uint_fast8_t count;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#define softfloat_countLeadingZeros64 softfloat_countLeadingZeros64
#include "primitives.h"
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros64( uint64_t a )
{
uint_fast8_t count;
@@ -46,6 +46,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| location pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
@@ -47,6 +47,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid exception
| is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
@@ -45,6 +45,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| location pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr )
{
@@ -45,6 +45,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| location pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_le128
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_le128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_lt128
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_lt128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_mul64ByShifted32To128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b )
{
uint_fast64_t mid;
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_mul64To128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64To128( uint64_t a, uint64_t b )
{
uint32_t a32, a0, b32, b0;
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
bool sign,
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig64 softfloat_normSubnormalExtF80Sig( uint_fast64_t sig )
{
int_fast8_t shiftDist;
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig128
softfloat_normSubnormalF128Sig( uint_fast64_t sig64, uint_fast64_t sig0 )
{
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig32 softfloat_normSubnormalF32Sig( uint_fast32_t sig )
{
int_fast8_t shiftDist;
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig64 softfloat_normSubnormalF64Sig( uint_fast64_t sig )
{
int_fast8_t shiftDist;
@@ -50,6 +50,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| result. If either original floating-point value is a signaling NaN, the
| invalid exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
uint_fast16_t uiA64,
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
bool sign,
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float32_t
softfloat_roundPackToF32( bool sign, int_fast16_t exp, uint_fast32_t sig )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float64_t
softfloat_roundPackToF64( bool sign, int_fast16_t exp, uint_fast64_t sig )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
softfloat_roundToI32(
bool sign, uint_fast64_t sig, uint_fast8_t roundingMode, bool exact )
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
bool sign,
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shiftRightJam32
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_shiftRightJam32( uint32_t a, uint_fast16_t dist )
{
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shiftRightJam64
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shiftRightJam64( uint64_t a, uint_fast32_t dist )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shiftRightJam64Extra
FEXCORE_PRESERVE_ALL_ATTR
struct uint64_extra
softfloat_shiftRightJam64Extra(
uint64_t a, uint64_t extra, uint_fast32_t dist )
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shortShiftLeft128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftLeft128( uint64_t a64, uint64_t a0, uint_fast8_t dist )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shortShiftRight128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftRight128( uint64_t a64, uint64_t a0, uint_fast8_t dist )
{
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shortShiftRightJam64
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shortShiftRightJam64( uint64_t a, uint_fast8_t dist )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_sub128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_sub128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -92,6 +92,7 @@ enum {
/*----------------------------------------------------------------------------
| Routine to raise any or all of the software floating-point exception flags.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t );
/*----------------------------------------------------------------------------
@@ -110,6 +111,7 @@ float16_t ui64_to_f16( uint64_t );
float32_t ui64_to_f32( uint64_t );
float64_t ui64_to_f64( uint64_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t ui64_to_extF80( uint64_t );
float128_t ui64_to_f128( uint64_t );
#endif
@@ -119,6 +121,7 @@ float16_t i32_to_f16( int32_t );
float32_t i32_to_f32( int32_t );
float64_t i32_to_f64( int32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t i32_to_extF80( int32_t );
float128_t i32_to_f128( int32_t );
#endif
@@ -183,6 +186,7 @@ int_fast64_t f32_to_i64_r_minMag( float32_t, bool );
float16_t f32_to_f16( float32_t );
float64_t f32_to_f64( float32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t );
float128_t f32_to_f128( float32_t );
#endif
@@ -218,6 +222,7 @@ int_fast64_t f64_to_i64_r_minMag( float64_t, bool );
float16_t f64_to_f16( float64_t );
float32_t f64_to_f32( float64_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t );
float128_t f64_to_f128( float64_t );
#endif
@@ -250,26 +255,41 @@ extern THREAD_LOCAL uint_fast8_t extF80_roundingPrecision;
*----------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_FAST_INT64
uint_fast32_t extF80_to_ui32( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t extF80_to_ui64( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t extF80_to_i32( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t extF80_to_i64( extFloat80_t, uint_fast8_t, bool );
uint_fast32_t extF80_to_ui32_r_minMag( extFloat80_t, bool );
uint_fast64_t extF80_to_ui64_r_minMag( extFloat80_t, bool );
int_fast32_t extF80_to_i32_r_minMag( extFloat80_t, bool );
int_fast64_t extF80_to_i64_r_minMag( extFloat80_t, bool );
float16_t extF80_to_f16( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_roundToInt( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t, extFloat80_t );
bool extF80_le( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t, extFloat80_t );
bool extF80_eq_signaling( extFloat80_t, extFloat80_t );
bool extF80_le_quiet( extFloat80_t, extFloat80_t );
@@ -320,6 +340,7 @@ int_fast64_t f128_to_i64_r_minMag( float128_t, bool );
float16_t f128_to_f16( float128_t );
float32_t f128_to_f32( float128_t );
float64_t f128_to_f64( float128_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t );
float128_t f128_roundToInt( float128_t, uint_fast8_t, bool );
float128_t f128_add( float128_t, float128_t );
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| to substitute a result value. If traps are not implemented, this routine
| should be simply `softfloat_exceptionFlags |= flags;'.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t flags )
{
@@ -135,12 +135,14 @@ uint_fast16_t
| location pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 32-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast32_t softfloat_commonNaNToF32UI( const struct commonNaN *aPtr );
/*----------------------------------------------------------------------------
@@ -170,12 +172,14 @@ uint_fast32_t
| location pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 64-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t softfloat_commonNaNToF64UI( const struct commonNaN *aPtr );
/*----------------------------------------------------------------------------
@@ -215,6 +219,7 @@ uint_fast64_t
| location pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
@@ -224,6 +229,7 @@ void
| floating-point NaN, and returns the bit pattern of this value as an unsigned
| integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr );
/*----------------------------------------------------------------------------
@@ -235,6 +241,7 @@ struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr );
| result. If either original floating-point value is a signaling NaN, the
| invalid exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
uint_fast16_t uiA64,
@@ -264,6 +271,7 @@ struct uint128
| pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid exception
| is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
@@ -272,6 +280,7 @@ void
| Converts the common NaN pointed to by 'aPtr' into a 128-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToF128UI( const struct commonNaN * );
/*----------------------------------------------------------------------------
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t ui64_to_extF80( uint64_t a )
{
uint_fast16_t uiZ64;
+20
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/BitUtils.h>
@@ -62,6 +63,7 @@ struct FEX_PACKED X80SoftFloat {
}
// Ops
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FADD(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -83,6 +85,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSUB(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -104,6 +107,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FMUL(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -125,6 +129,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FDIV(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -146,6 +151,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FREM(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -168,6 +174,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FREM1(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -190,14 +197,17 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -221,6 +231,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FXTRACT_EXP(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -242,12 +253,14 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static void FCMP(X80SoftFloat const &lhs, X80SoftFloat const &rhs, bool *eq, bool *lt, bool *nan) {
*eq = extF80_eq(lhs, rhs);
*lt = extF80_lt(lhs, rhs);
*nan = IsNan(lhs) || IsNan(rhs);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSCALE(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -276,6 +289,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat F2XM1(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -299,6 +313,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FYL2X(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -324,6 +339,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FATAN(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -349,6 +365,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FTAN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -372,6 +389,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSIN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -394,6 +412,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FCOS(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -416,6 +435,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSQRT(X80SoftFloat const &lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
+1
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/string.h>
+1
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/string.h>
+3 -12
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include "Common/StringConv.h"
#include "Common/StringUtils.h"
#include "FEXCore/Utils/EnumUtils.h"
@@ -334,16 +335,11 @@ namespace DefaultValues {
// Sanitize Core option
FEX_CONFIG_OPT(Core, CORE);
#if (_M_X86_64)
constexpr uint32_t MaxCoreNumber = 2;
#else
constexpr uint32_t MaxCoreNumber = 1;
#endif
#ifdef INTERPRETER_ENABLED
constexpr uint32_t MinCoreNumber = 0;
#else
constexpr uint32_t MinCoreNumber = 1;
constexpr uint32_t MaxCoreNumber = 0;
#endif
if (Core > MaxCoreNumber || Core < MinCoreNumber) {
if (Core > MaxCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, fextl::fmt::format("{}", static_cast<uint32_t>(FEXCore::Config::CONFIG_IRJIT)));
}
@@ -352,11 +348,6 @@ namespace DefaultValues {
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION)) {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(Core, CORE);
if (CacheObjectCodeCompilation() && Core() == FEXCore::Config::CONFIG_INTERPRETER) {
// If running the interpreter then disable cache code compilation
FEXCore::Config::Erase(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION);
}
}
fextl::string ContainerPrefix { FindContainerPrefix() };
+22 -4
View File
@@ -6,12 +6,12 @@
"Default": "FEXCore::Config::ConfigCore::CONFIG_IRJIT",
"TextDefault": "irjit",
"ShortArg": "c",
"Choices": [ "irint", "irjit", "host" ],
"Choices": [ "irjit", "host" ],
"ArgumentHandler": "CoreHandler",
"Desc": [
"Which CPU core to use",
"host only exists on x86_64",
"[irint, irjit, host]"
"[irjit, host]"
]
},
"Multiblock": {
@@ -59,6 +59,8 @@
"DISABLESVE": "disablesve",
"ENABLEAVX": "enableavx",
"DISABLEAVX": "disableavx",
"ENABLEAVX2": "enableavx2",
"DISABLEAVX2": "disableavx2",
"ENABLEAFP": "enableafp",
"DISABLEAFP": "disableafp",
"ENABLELRCPC": "enablelrcpc",
@@ -76,13 +78,18 @@
"ENABLEATOMICS": "enableatomics",
"DISABLEATOMICS": "disableatomics",
"ENABLEFCMA": "enablefcma",
"DISABLEFCMA": "disablefcma"
"DISABLEFCMA": "disablefcma",
"ENABLEFLAGM": "enableflagm",
"DISABLEFLAGM": "disableflagm",
"ENABLEFLAGM2": "enableflagm2",
"DISABLEFLAGM2": "disableflagm2"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
"\toff: Default CPU features queried from CPU features",
"\t{enable,disable}sve: Will force enable or disable sve even if the host doesn't support it",
"\t{enable,disable}avx: Will force enable or disable avx even if the host doesn't support it",
"\t{enable,disable}avx2: Will force enable or disable avx2 even if the host doesn't support it",
"\t{enable,disable}afp: Will force enable or disable afp even if the host doesn't support it",
"\t{enable,disable}lrcpc: Will force enable or disable lrcpc even if the host doesn't support it",
"\t{enable,disable}lrcpc2: Will force enable or disable lrcpc2 even if the host doesn't support it",
@@ -91,7 +98,9 @@
"\t{enable,disable}rng: Will force enable or disable rng even if the host doesn't support it",
"\t{enable,disable}clzero: Will force enable or disable clzero even if the host doesn't support it",
"\t{enable,disable}atomics: Will force enable or disable ARMv8.1 LSE atomics even if the host doesn't support it",
"\t{enable,disable}fcma: Will force enable or disable fcma even if the host doesn't support it"
"\t{enable,disable}fcma: Will force enable or disable fcma even if the host doesn't support it",
"\t{enable,disable}flagm: Will force enable or disable flagm even if the host doesn't support it",
"\t{enable,disable}flagm2: Will force enable or disable flagm2 even if the host doesn't support it"
]
}
},
@@ -490,6 +499,15 @@
"IS64BIT_MODE": {
"Type": "bool",
"Default": "false"
},
"DISABLE_VIXL_INDIRECT_RUNTIME_CALLS": {
"Type": "bool",
"Default": "true",
"Desc": [
"This option is used for the InstructionCountCI so it can generate the same codegen between Arm64 hosts and vixl simulator hosts.",
"Vixl simulator indirect runtime calls are a special hlt instruction with metadata after it. Effectively making a custom call instruction.",
"With visual simulator calls disabled, the code generation would be the same as on a native Arm64 host, but running the code is broken."
]
}
}
}
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
+18 -2
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Common/JitSymbols.h"
@@ -45,7 +46,6 @@ namespace CodeSerialize {
namespace CPU {
class Arm64JITCore;
class X86JITCore;
class InterpreterCore;
class Dispatcher;
}
namespace HLE {
@@ -87,6 +87,8 @@ namespace FEXCore::Context {
ExitReason RunUntilExit() override;
void ExecuteThread(FEXCore::Core::InternalThreadState *Thread) override;
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) override;
int GetProgramStatus() const override;
@@ -204,7 +206,6 @@ namespace FEXCore::Context {
friend class FEXCore::CPU::X86JITCore;
#endif
friend class FEXCore::CPU::InterpreterCore;
friend class FEXCore::IR::Validation::IRValidation;
struct {
@@ -240,6 +241,7 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
} Config;
FEXCore::HostFeatures HostFeatures;
@@ -378,6 +380,20 @@ namespace FEXCore::Context {
UpdateAtomicTSOEmulationConfig();
}
// Returns if Software TSO emulation is required.
// NOTE: This doesn't necessary return if Atomic-based TSO is currently enabled.
// This will still return true if on a single thread and TSO is currently disabled.
//
// This is to ensure that if early initialization checks CPU features and TSO /could/ be enabled, that
// we return consistent results.
//
// To check if Atomic TSO is currently enabled in the JIT, use `IsAtomicTSOEnabled` instead.
bool SoftwareTSORequired() const {
if (SupportsHardwareTSO) return false;
return Config.TSOEnabled;
}
void EnableExitOnHLT() override { ExitOnHLT = true; }
bool ExitOnHLTEnabled() const { return ExitOnHLT; }
@@ -1,6 +1,8 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include "Interface/HLE/Thunks/Thunks.h"
@@ -66,11 +68,102 @@ namespace x64 {
};
// v8..v15 = (lower 64bits) Callee saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 12> RAFPR = {
// v0 ~ v3 are used as temps.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> RAFPR = {
// v0 ~ v1 are used as temps.
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
// FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
};
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 7> PreserveAll_SRA = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8,
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
};
constexpr uint32_t PreserveAll_SRAMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRA) {
switch (Reg.Idx()) {
case 0:
case 1:
case 2:
case 3:
case 4:
case 5:
case 6:
case 7:
case 8:
case 16:
case 17:
Mask |= (1U << Reg.Idx());
break;
default: break;
}
}
return Mask;
}()
};
// Dynamic GPRs
constexpr std::array<FEXCore::ARMEmitter::Register, 1> PreserveAll_Dynamic = {
// Only LR needs to get saved.
FEXCore::ARMEmitter::Reg::r30
};
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
// None.
};
constexpr uint32_t PreserveAll_SRAFPRMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPR) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs
// - v0-v7
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
// v0 ~ v1 are temps
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
};
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
// This is /all/ of the SRA registers
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPRSVE) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs when the host supports SVE-256bit.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> PreserveAll_DynamicFPRSVE = {
// v0 ~ v1 are used as temps.
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
@@ -127,11 +220,106 @@ namespace x32 {
};
// v8..v15 = (lower 64bits) Callee saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 20> RAFPR = {
// v0 ~ v3 are used as temps.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> RAFPR = {
// v0 ~ v1 are used as temps.
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
// FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25,
FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29,
FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31
};
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 5> PreserveAll_SRA = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8,
};
constexpr uint32_t PreserveAll_SRAMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRA) {
switch (Reg.Idx()) {
case 0:
case 1:
case 2:
case 3:
case 4:
case 5:
case 6:
case 7:
case 8:
case 16:
case 17:
Mask |= (1U << Reg.Idx());
break;
default: break;
}
}
return Mask;
}()
};
// Dynamic GPRs
constexpr std::array<FEXCore::ARMEmitter::Register, 3> PreserveAll_Dynamic = {
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
FEXCore::ARMEmitter::Reg::r30
};
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
// None.
};
constexpr uint32_t PreserveAll_SRAFPRMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPR) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs
// - v0-v7
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
// v0 ~ v1 are temps
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
};
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
// This is /all/ of the SRA registers
constexpr std::array<FEXCore::ARMEmitter::VRegister, 8> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPRSVE) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs when the host supports SVE-256bit.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> PreserveAll_DynamicFPRSVE = {
// v0 ~ v1 are used as temps.
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
@@ -532,6 +720,107 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
}
}
void Arm64Emitter::PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
if (SVERegs) {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
st2b(Reg1.Z(), Reg2.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 2);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
st4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 4);
}
}
else {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), TmpReg, 32);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), TmpReg, 64);
}
}
}
void Arm64Emitter::PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs) {
size_t i = 0;
for (; i < (Regs.size() % 2); ++i) {
const auto Reg1 = Regs[i];
str<ARMEmitter::IndexType::POST>(Reg1.X(), TmpReg, 16);
}
for (; i < Regs.size(); i += 2) {
const auto Reg1 = Regs[i];
const auto Reg2 = Regs[i + 1];
stp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), TmpReg, 16);
}
}
void Arm64Emitter::PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
if (SVERegs) {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
ld2b(Reg1.Z(), Reg2.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 2);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
ld4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 4);
}
} else {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), ARMEmitter::Reg::rsp, 32);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), ARMEmitter::Reg::rsp, 64);
}
}
}
void Arm64Emitter::PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs) {
size_t i = 0;
for (; i < (Regs.size() % 2); ++i) {
const auto Reg1 = Regs[i];
ldr<ARMEmitter::IndexType::POST>(Reg1.X(), ARMEmitter::Reg::rsp, 16);
}
for (; i < Regs.size(); i += 2) {
const auto Reg1 = Regs[i];
const auto Reg2 = Regs[i + 1];
ldp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), ARMEmitter::Reg::rsp, 16);
}
}
void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto GPRSize = (ConfiguredDynamicRegisterBase.size() + 1) * Core::CPUState::GPR_REG_SIZE;
@@ -545,31 +834,13 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
// rsp capable move
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
if (CanUseSVE) {
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
st4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 4);
}
} else {
LOGMAN_THROW_A_FMT(GeneralFPRegisters.size() % 4 == 0, "Needs to have multiple of 4 FPRs for RA");
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), TmpReg, 64);
}
}
LOGMAN_THROW_A_FMT(GeneralFPRegisters.size() % 2 == 0, "Needs to have multiple of 2 FPRs for RA");
for (size_t i = 0; i < ConfiguredDynamicRegisterBase.size(); i += 2) {
const auto Reg1 = ConfiguredDynamicRegisterBase[i];
const auto Reg2 = ConfiguredDynamicRegisterBase[i + 1];
stp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), TmpReg, 16);
}
// Push the vector registers
PushVectorRegisters(TmpReg, CanUseSVE, GeneralFPRegisters);
// Push the general registers.
PushGeneralRegisters(TmpReg, ConfiguredDynamicRegisterBase);
str(ARMEmitter::XReg::lr, TmpReg, 0);
}
@@ -577,34 +848,107 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
void Arm64Emitter::PopDynamicRegsAndLR() {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
if (CanUseSVE) {
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
ld4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 4);
}
} else {
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), ARMEmitter::Reg::rsp, 64);
}
}
// Pop vectors first
PopVectorRegisters(CanUseSVE, GeneralFPRegisters);
for (size_t i = 0; i < ConfiguredDynamicRegisterBase.size(); i += 2) {
const auto Reg1 = ConfiguredDynamicRegisterBase[i];
const auto Reg2 = ConfiguredDynamicRegisterBase[i + 1];
ldp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), ARMEmitter::Reg::rsp, 16);
}
// Pop GPRs second
PopGeneralRegisters(ConfiguredDynamicRegisterBase);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
}
void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs{};
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs{};
uint32_t PreserveSRAMask{};
uint32_t PreserveSRAFPRMask{};
if (EmitterCTX->Config.Is64BitMode()) {
DynamicGPRs = x64::PreserveAll_Dynamic;
DynamicFPRs = x64::PreserveAll_DynamicFPR;
PreserveSRAMask = x64::PreserveAll_SRAMask;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
}
}
else {
DynamicGPRs = x32::PreserveAll_Dynamic;
DynamicFPRs = x32::PreserveAll_DynamicFPR;
PreserveSRAMask = x32::PreserveAll_SRAMask;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
}
}
const auto GPRSize = AlignUp(DynamicGPRs.size(), 2) * Core::CPUState::GPR_REG_SIZE;
const auto FPRSize = DynamicFPRs.size() * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
// Spill the static registers.
SpillStaticRegs(TmpReg, true, PreserveSRAMask, PreserveSRAFPRMask);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
// rsp capable move
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
// Push the vector registers.
PushVectorRegisters(TmpReg, CanUseSVE, DynamicFPRs);
// Push the general registers.
PushGeneralRegisters(TmpReg, DynamicGPRs);
}
void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs{};
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs{};
uint32_t PreserveSRAMask{};
uint32_t PreserveSRAFPRMask{};
if (EmitterCTX->Config.Is64BitMode()) {
DynamicGPRs = x64::PreserveAll_Dynamic;
DynamicFPRs = x64::PreserveAll_DynamicFPR;
PreserveSRAMask = x64::PreserveAll_SRAMask;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
}
}
else {
DynamicGPRs = x32::PreserveAll_Dynamic;
DynamicFPRs = x32::PreserveAll_DynamicFPR;
PreserveSRAMask = x32::PreserveAll_SRAMask;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
}
}
// Fill the static registers.
FillStaticRegs(true, PreserveSRAMask, PreserveSRAFPRMask);
// Pop the vector registers.
PopVectorRegisters(CanUseSVE, DynamicFPRs);
// Pop the general registers.
PopGeneralRegisters(DynamicGPRs);
}
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetCursorAddress<uint64_t>();
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
@@ -1,10 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include <aarch64/assembler-aarch64.h>
@@ -28,6 +28,10 @@
#include <utility>
#include <span>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
constexpr auto STATE = FEXCore::ARMEmitter::XReg::x28;
@@ -42,8 +46,6 @@ constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x3;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v0;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
constexpr auto VTMP3 = FEXCore::ARMEmitter::VReg::v2;
constexpr auto VTMP4 = FEXCore::ARMEmitter::VReg::v3;
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
@@ -99,12 +101,52 @@ protected:
// We can't guarantee only the lower 64bits are used so flush everything
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
// Generic push and pop vector registers.
void PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs);
void PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs);
void PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg);
void PopDynamicRegsAndLR();
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
// Spills and fills SRA/Dynamic registers that are required for Arm64 `preserve_all` ABI.
// This ABI changes most registers to be callee saved.
// Caller Saved:
// - X0-X8, X16-X18.
// - v0-v7
// - For 256-bit SVE hosts: top 128-bits of v8-v31
//
// Callee Saved:
// - X9-X15, X19-X31
// - Low 128-bits of v8-v31
void SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true);
void FillForPreserveAllABICall(bool FPRs = true);
void SpillForABICall(bool SupportsPreserveAllABI, FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true) {
if (SupportsPreserveAllABI) {
SpillForPreserveAllABICall(TMP1, true);
}
else {
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
}
}
void FillForABICall(bool SupportsPreserveAllABI, bool FPRs = true) {
if (SupportsPreserveAllABI) {
FillForPreserveAllABICall(true);
}
else {
PopDynamicRegsAndLR();
FillStaticRegs();
}
}
void Align16B();
#ifdef VIXL_SIMULATOR
@@ -171,7 +213,15 @@ protected:
// Call type
dc32(vixl::aarch64::kCallRuntime);
}
#else
template<typename R, typename... P>
void GenerateRuntimeCall(R (*Function)(P...)) {
// Explicitly doing nothing.
}
template<typename R, typename... P>
void GenerateIndirectRuntimeCall(ARMEmitter::Register Reg) {
// Explicitly doing nothing.
}
#endif
#ifdef VIXL_SIMULATOR
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* ALU instruction emitters.
*
* Almost all of these operations have `ARMEmitter::Size` as their first argument.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* ASIMD instruction emitters.
*
* This contains emitters for vector operations explicitly.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* Branch instruction emitters.
*
* Most of these instructions will use `BackwardLabel`, `ForwardLabel`, or `BiDirectionLabel` to determine where a branch targets.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef>
#include <cstdint>
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* Load-store instruction emitters
*
* For GPR load-stores that take a `Size` argument as their first argument can be 32-bit or 64-bit.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/EnumUtils.h>
@@ -20,12 +21,11 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const Register&, const Register&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
WRegister W() const;
XRegister X() const;
constexpr WRegister W() const;
constexpr XRegister X() const;
private:
uint32_t Index;
@@ -45,16 +45,15 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const WRegister&, const WRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator Register() const {
constexpr operator Register() const {
return Register(Index);
}
XRegister X() const;
Register R() const;
constexpr XRegister X() const;
constexpr Register R() const;
private:
uint32_t Index;
@@ -74,16 +73,15 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const XRegister&, const XRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator Register() const {
constexpr operator Register() const {
return Register(Index);
}
WRegister W() const;
Register R() const;
constexpr WRegister W() const;
constexpr Register R() const;
private:
uint32_t Index;
@@ -92,27 +90,27 @@ namespace FEXCore::ARMEmitter {
static_assert(std::is_trivial_v<Register>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<Register>, "Needs to be standard");
inline WRegister Register::W() const {
inline constexpr WRegister Register::W() const {
return WRegister{Index};
}
inline XRegister Register::X() const {
inline constexpr XRegister Register::X() const {
return XRegister{Index};
}
inline XRegister WRegister::X() const {
inline constexpr XRegister WRegister::X() const {
return XRegister{Index};
}
inline Register WRegister::R() const {
inline constexpr Register WRegister::R() const {
return *this;
}
inline WRegister XRegister::W() const {
inline constexpr WRegister XRegister::W() const {
return WRegister{Index};
}
inline Register XRegister::R() const {
inline constexpr Register XRegister::R() const {
return *this;
}
@@ -259,7 +257,6 @@ namespace FEXCore::ARMEmitter {
class QRegister;
class ZRegister;
/* Unsized ASIMD register class
* This class doesn't imply a size when used, nor implies Vector or Scalar.
* It does imply that this instruction isn't using the register for SVE.
@@ -272,16 +269,16 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const VRegister&, const VRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
BRegister B() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -301,20 +298,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const BRegister&, const BRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister () const {
constexpr operator VRegister() const {
return VRegister(Index);
}
BRegister V() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr BRegister V() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -334,20 +330,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const HRegister&, const HRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister() const {
constexpr operator VRegister() const {
return VRegister(Index);
}
HRegister V() const;
BRegister B() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr HRegister V() const;
constexpr BRegister B() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -367,20 +362,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const SRegister&, const SRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister() const {
constexpr operator VRegister() const {
return VRegister(Index);
}
SRegister V() const;
BRegister B() const;
HRegister H() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr SRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -401,20 +395,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const DRegister&, const DRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister() const {
constexpr operator VRegister() const {
return VRegister(Index);
}
DRegister V() const;
BRegister B() const;
HRegister H() const;
SRegister S() const;
QRegister Q() const;
ZRegister Z() const;
constexpr DRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -435,20 +428,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const QRegister&, const QRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister () const {
constexpr operator VRegister() const {
return VRegister(Index);
}
QRegister V() const;
BRegister B() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
ZRegister Z() const;
constexpr QRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -468,16 +460,16 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const ZRegister&, const ZRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
VRegister V() const;
BRegister B() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
constexpr VRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
private:
uint32_t Index;
@@ -487,142 +479,142 @@ namespace FEXCore::ARMEmitter {
static_assert(std::is_standard_layout_v<ZRegister>, "Needs to be standard");
// VRegister
inline BRegister VRegister::B() const {
inline constexpr BRegister VRegister::B() const {
return BRegister{Index};
}
inline HRegister VRegister::H() const {
inline constexpr HRegister VRegister::H() const {
return HRegister{Index};
}
inline SRegister VRegister::S() const {
inline constexpr SRegister VRegister::S() const {
return SRegister{Index};
}
inline DRegister VRegister::D() const {
inline constexpr DRegister VRegister::D() const {
return DRegister{Index};
}
inline QRegister VRegister::Q() const {
inline constexpr QRegister VRegister::Q() const {
return QRegister{Index};
}
inline ZRegister VRegister::Z() const {
inline constexpr ZRegister VRegister::Z() const {
return ZRegister{Index};
}
// BRegister
inline BRegister BRegister::V() const {
inline constexpr BRegister BRegister::V() const {
return *this;
}
inline HRegister BRegister::H() const {
inline constexpr HRegister BRegister::H() const {
return HRegister{Index};
}
inline SRegister BRegister::S() const {
inline constexpr SRegister BRegister::S() const {
return SRegister{Index};
}
inline DRegister BRegister::D() const {
inline constexpr DRegister BRegister::D() const {
return DRegister{Index};
}
inline QRegister BRegister::Q() const {
inline constexpr QRegister BRegister::Q() const {
return QRegister{Index};
}
inline ZRegister BRegister::Z() const {
inline constexpr ZRegister BRegister::Z() const {
return ZRegister{Index};
}
// HRegister
inline HRegister HRegister::V() const {
inline constexpr HRegister HRegister::V() const {
return *this;
}
inline BRegister HRegister::B() const {
inline constexpr BRegister HRegister::B() const {
return BRegister{Index};
}
inline SRegister HRegister::S() const {
inline constexpr SRegister HRegister::S() const {
return SRegister{Index};
}
inline DRegister HRegister::D() const {
inline constexpr DRegister HRegister::D() const {
return DRegister{Index};
}
inline QRegister HRegister::Q() const {
inline constexpr QRegister HRegister::Q() const {
return QRegister{Index};
}
inline ZRegister HRegister::Z() const {
inline constexpr ZRegister HRegister::Z() const {
return ZRegister{Index};
}
// SRegister
inline SRegister SRegister::V() const {
inline constexpr SRegister SRegister::V() const {
return *this;
}
inline BRegister SRegister::B() const {
inline constexpr BRegister SRegister::B() const {
return BRegister{Index};
}
inline HRegister SRegister::H() const {
inline constexpr HRegister SRegister::H() const {
return HRegister{Index};
}
inline DRegister SRegister::D() const {
inline constexpr DRegister SRegister::D() const {
return DRegister{Index};
}
inline QRegister SRegister::Q() const {
inline constexpr QRegister SRegister::Q() const {
return QRegister{Index};
}
inline ZRegister SRegister::Z() const {
inline constexpr ZRegister SRegister::Z() const {
return ZRegister{Index};
}
// DRegister
inline DRegister DRegister::V() const {
inline constexpr DRegister DRegister::V() const {
return DRegister{Index};
}
inline BRegister DRegister::B() const {
inline constexpr BRegister DRegister::B() const {
return BRegister{Index};
}
inline HRegister DRegister::H() const {
inline constexpr HRegister DRegister::H() const {
return HRegister{Index};
}
inline SRegister DRegister::S() const {
inline constexpr SRegister DRegister::S() const {
return SRegister{Index};
}
inline QRegister DRegister::Q() const {
inline constexpr QRegister DRegister::Q() const {
return QRegister{Index};
}
inline ZRegister DRegister::Z() const {
inline constexpr ZRegister DRegister::Z() const {
return ZRegister{Index};
}
// QRegister
inline QRegister QRegister::V() const {
inline constexpr QRegister QRegister::V() const {
return *this;
}
inline BRegister QRegister::B() const {
inline constexpr BRegister QRegister::B() const {
return BRegister{Index};
}
inline HRegister QRegister::H() const {
inline constexpr HRegister QRegister::H() const {
return HRegister{Index};
}
inline SRegister QRegister::S() const {
inline constexpr SRegister QRegister::S() const {
return SRegister{Index};
}
inline DRegister QRegister::D() const {
inline constexpr DRegister QRegister::D() const {
return DRegister{Index};
}
inline ZRegister QRegister::Z() const {
inline constexpr ZRegister QRegister::Z() const {
return ZRegister{Index};
}
// ZRegister
inline VRegister ZRegister::V() const {
inline constexpr VRegister ZRegister::V() const {
return VRegister(Index);
}
inline BRegister ZRegister::B() const {
inline constexpr BRegister ZRegister::B() const {
return BRegister(Index);
}
inline HRegister ZRegister::H() const {
inline constexpr HRegister ZRegister::H() const {
return HRegister(Index);
}
inline SRegister ZRegister::S() const {
inline constexpr SRegister ZRegister::S() const {
return SRegister(Index);
}
inline DRegister ZRegister::D() const {
inline constexpr DRegister ZRegister::D() const {
return DRegister(Index);
}
inline QRegister ZRegister::Q() const {
inline constexpr QRegister ZRegister::Q() const {
return QRegister(Index);
}
@@ -879,36 +871,28 @@ namespace FEXCore::ARMEmitter {
}
// Zero-cost FPR->GPR
inline
Register ToReg(HRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(HRegister Reg) {
return Register(Reg.Idx());
}
inline
Register ToReg(SRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(SRegister Reg) {
return Register(Reg.Idx());
}
inline
Register ToReg(DRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(DRegister Reg) {
return Register(Reg.Idx());
}
inline
Register ToReg(VRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(VRegister Reg) {
return Register(Reg.Idx());
}
// Zero-cost GPR->FPR
inline
VRegister ToVReg(Register Reg) {
return static_cast<VRegister>(Reg.Idx());
inline constexpr VRegister ToVReg(Register Reg) {
return VRegister(Reg.Idx());
}
inline
VRegister ToVReg(XRegister Reg) {
return static_cast<VRegister>(Reg.Idx());
inline constexpr VRegister ToVReg(XRegister Reg) {
return VRegister(Reg.Idx());
}
inline
VRegister ToVReg(WRegister Reg) {
return static_cast<VRegister>(Reg.Idx());
inline constexpr VRegister ToVReg(WRegister Reg) {
return VRegister(Reg.Idx());
}
class PRegisterZero;
@@ -925,12 +909,12 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const PRegister&, const PRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
PRegisterZero Zeroing() const;
PRegisterMerge Merging() const;
constexpr PRegisterZero Zeroing() const;
constexpr PRegisterMerge Merging() const;
private:
uint32_t Index;
@@ -948,14 +932,17 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const PRegisterZero&, const PRegisterZero&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator PRegister() const;
PRegister P() const;
PRegisterMerge Merging() const;
constexpr operator PRegister() const {
return PRegister(Index);
}
constexpr PRegister P() const {
return PRegister(Index);
}
constexpr PRegisterMerge Merging() const;
private:
uint32_t Index;
@@ -973,14 +960,17 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const PRegisterMerge&, const PRegisterMerge&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator PRegister() const;
PRegister P() const;
PRegisterZero Zeroing() const;
constexpr operator PRegister() const {
return PRegister(Index);
}
constexpr PRegister P() const {
return PRegister(Index);
}
constexpr PRegisterZero Zeroing() const;
private:
uint32_t Index;
@@ -989,39 +979,21 @@ namespace FEXCore::ARMEmitter {
static_assert(std::is_trivial_v<PRegisterZero>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<PRegisterZero>, "Needs to be standard");
// PRegister
inline PRegisterZero PRegister::Zeroing() const {
inline constexpr PRegisterZero PRegister::Zeroing() const {
return PRegisterZero(Idx());
}
inline PRegisterMerge PRegister::Merging() const {
inline constexpr PRegisterMerge PRegister::Merging() const {
return PRegisterMerge(Idx());
}
// PRegisterZero
inline PRegisterZero::operator PRegister() const {
return PRegister(Index);
}
inline PRegister PRegisterZero::P() const {
return PRegister(Idx());
}
inline PRegisterMerge PRegisterZero::Merging() const {
inline constexpr PRegisterMerge PRegisterZero::Merging() const {
return PRegisterMerge(Idx());
}
// PRegisterMerge
inline PRegisterMerge::operator PRegister() const {
return PRegisterZero(Index);
}
inline PRegister PRegisterMerge::P() const {
return PRegister(Idx());
}
inline PRegisterZero PRegisterMerge::Zeroing() const {
inline constexpr PRegisterZero PRegisterMerge::Zeroing() const {
return PRegisterZero(Idx());
}
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* SVE instruction emitters
* These contain instruction emitters for AArch64 SVE and SVE2 operations.
*
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* Scalar instruction emitters.
*
* These contain instruction emitters for scalar ASIMD operations explicitly.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* System instruction emitters.
*
* This is mostly a mashup of various instruction types.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/BlockSamplingData.h"
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstdint>
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
@@ -126,6 +127,55 @@ constexpr static auto PSHUFD_LUT {
}()
};
constexpr static auto SHUFPS_LUT {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// 32-bit words in [127:96], [95:64], [63:32], [31:0] are selected using the 8-bit Index.
// Expectation for this LUT is to simulate SHUFPS with ARM's TBL (two register) instruction.
// SHUFPS behaviour:
// Two 32-bits words from each source are selected from each source in the lower and upper halves of the 128-bit destination.
// Dest[31:0] = Src1[<Word0>]
// Dest[63:32] = Src1[<Word1>]
// Dest[95:64] = Src2[<Word2>]
// Dest[127:96] = Src2[<Word3>]
std::array<LUTType, 256> TotalLUT{};
const uint64_t WordSelectionSrc1[4] = {
0x03'02'01'00,
0x07'06'05'04,
0x0b'0a'09'08,
0x0f'0e'0d'0c,
};
// Src2 needs to offset each byte index by 16-bytes to pull from the second source.
const uint64_t WordSelectionSrc2[4] = {
0x03'02'01'00 + (0x10101010),
0x07'06'05'04 + (0x10101010),
0x0b'0a'09'08 + (0x10101010),
0x0f'0e'0d'0c + (0x10101010),
};
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] =
(WordSelectionSrc1[Word0] << 0) |
(WordSelectionSrc1[Word1] << 32);
LUT.Val[1] =
(WordSelectionSrc2[Word2] << 0) |
(WordSelectionSrc2[Word3] << 32);
}
return TotalLUT;
}()
};
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState), InitialCodeSize(InitialCodeSize), MaxCodeSize(MaxCodeSize) {
@@ -136,10 +186,14 @@ CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t I
Common.NamedVectorConstantPointers[i] = reinterpret_cast<uint64_t>(NamedVectorConstants[i]);
}
// Copy named vector constants.
memcpy(Common.NamedVectorConstants, NamedVectorConstants, sizeof(NamedVectorConstants));
// Initialize Indexed named vector constants.
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFLW] = reinterpret_cast<uint64_t>(PSHUFLW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFHW] = reinterpret_cast<uint64_t>(PSHUFHW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFD] = reinterpret_cast<uint64_t>(PSHUFD_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_SHUFPS] = reinterpret_cast<uint64_t>(SHUFPS_LUT.data());
#ifndef FEX_DISABLE_TELEMETRY
// Fill in telemetry values
+34 -55
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: opcodes|cpuid
@@ -20,9 +21,6 @@ $end_info$
#include "git_version.h"
#include <cstring>
#ifdef _M_X86_64
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#endif
namespace FEXCore {
namespace ProductNames {
@@ -66,7 +64,6 @@ namespace ProductNames {
static const char ARM_Firestorm[] = "Apple Firestorm";
static const char ARM_Icestorm[] = "Apple Icestorm";
#else
static const char UNKNOWN[] = "Unknown CPU";
#endif
}
@@ -342,38 +339,15 @@ void CPUIDEmu::SetupHostHybridFlag() {
#else
static uint32_t GetCycleCounterFrequency() {
uint32_t data[4];
Xbyak::util::Cpu::getCpuid(0, data);
if (data[0] >= 0x15) {
Xbyak::util::Cpu::getCpuid(0x15, data);
if (data[0] && data[1] && data[2]) {
return data[2] * data[1] / data[0];
}
}
return 0;
}
void CPUIDEmu::SetupHostHybridFlag() {
uint32_t data[4];
Xbyak::util::Cpu::getCpuid(0, data);
if (data[0] >= 0x7) {
Xbyak::util::Cpu::getCpuid(0x7, data);
// Bit 15 of edx claims hybrid CPU
Hybrid = (data[3] & (1U << 15)) != 0;
}
size_t CPUs = FEXCore::CPUInfo::CalculateNumberOfCPUs();
PerCPUData.resize(CPUs);
for (size_t i = 0; i < CPUs; ++i) {
PerCPUData[i].IsBig = true;
PerCPUData[i].ProductName = ProductNames::UNKNOWN;
}
}
#endif
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
// EBX, EDX, ECX become the manufacturer id string
@@ -392,7 +366,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) {
}
// Processor Info and Features bits
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
uint32_t CoreCount = Cores();
@@ -477,7 +451,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
}
// 2: Cache and TLB information
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_02h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_02h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
// returns default values from i7 model 1Ah
@@ -502,7 +476,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_02h(uint32_t Leaf) {
}
// 4: Deterministic cache parameters for each level
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_04h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_04h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
constexpr uint32_t CacheType_Data = 1;
constexpr uint32_t CacheType_Instruction = 2;
@@ -608,16 +582,21 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_04h(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_06h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_06h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
Res.eax = (1 << 2); // Always running APIC
Res.ecx = (0 << 3); // Intel performance energy bias preference (EPB)
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
if (Leaf == 0) {
// Disable Enhanced REP MOVS when TSO is enabled.
// vcruntime140 memmove will use `rep movsb` in this case which completely destroys perf in Hades(appId 1145360)
// This is due to LRCPC performance on Cortex being abysmal.
const uint32_t SupportsEnhancedREPMOVS = CTX->SoftwareTSORequired() ? 0 : 1;
// Number of subfunctions
Res.eax = 0x0;
Res.ebx =
@@ -630,7 +609,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(1 << 6) | // FPU data pointer updated only on exception
(1 << 7) | // SMEP support
(SupportsAVX() << 8) | // BMI2
(0 << 9) | // Enhanced REP MOVSB/STOSB
(SupportsEnhancedREPMOVS << 9) | // Enhanced REP MOVSB/STOSB
(1 << 10) | // INVPCID for system software control of process-context
(0 << 11) | // Restricted transactional memory
(0 << 12) | // Intel resource directory technology Monitoring
@@ -726,7 +705,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) const {
// Leaf 0
FEXCore::CPUID::FunctionResults Res{};
@@ -780,7 +759,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
// TSC frequency = ECX * EBX / EAX
uint32_t FrequencyHz = GetCycleCounterFrequency();
@@ -792,7 +771,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
if (Hybrid) {
uint32_t CPU = GetCPUID();
@@ -805,7 +784,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) {
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
// Maximum supported hypervisor leafs
// We only expose the information leaf
@@ -827,7 +806,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
if (Leaf == 0) {
// EAX[3:0] Is the host architecture that FEX is running under
@@ -846,7 +825,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) {
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
Res.eax = 0x8000001F;
@@ -865,7 +844,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
}
// Extended processor and feature bits
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) const {
// RDTSCP is disabled on WIN32/Wine because there is no sane way to query processor ID.
#ifndef _WIN32
@@ -955,33 +934,33 @@ constexpr ssize_t DESCRIBE_STR_SIZE = std::char_traits<char>::length(GIT_DESCRIB
static_assert(DESCRIBE_STR_SIZE < 32);
//Processor brand string
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf) const {
return Function_8000_0002h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf) const {
return Function_8000_0003h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf) const {
return Function_8000_0004h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf, uint32_t CPU) const {
FEXCore::CPUID::FunctionResults Res{};
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[0], std::min(ssize_t{16L}, DESCRIBE_STR_SIZE));
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf, uint32_t CPU) const {
FEXCore::CPUID::FunctionResults Res{};
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[16], std::max(ssize_t{0L}, DESCRIBE_STR_SIZE - 16));
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf, uint32_t CPU) const {
FEXCore::CPUID::FunctionResults Res{};
auto &Data = PerCPUData[CPU];
memcpy(&Res, Data.ProductName, std::min(strlen(Data.ProductName), sizeof(FEXCore::CPUID::FunctionResults)));
@@ -989,7 +968,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf, uin
}
// L1 Cache and TLB identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0005h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0005h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
// L1 TLB Information for 2MB and 4MB pages
@@ -1024,7 +1003,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0005h(uint32_t Leaf) {
}
// L2 Cache identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0006h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0006h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
// L2 TLB Information for 2MB and 4MB pages
@@ -1058,7 +1037,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0006h(uint32_t Leaf) {
}
// Advanced power management
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0007h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0007h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
Res.eax = (1 << 2); // APIC timer not affected by p-state
Res.edx =
@@ -1067,7 +1046,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0007h(uint32_t Leaf) {
}
// Virtual and physical address sizes
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
Res.eax =
(48 << 0) | // PhysAddrSize = 48-bit
@@ -1089,7 +1068,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) {
}
// TLB 1GB page identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0019h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0019h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
Res.eax =
(0xF << 28) | // L1 DTLB associativity for 1GB pages
@@ -1106,7 +1085,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0019h(uint32_t Leaf) {
}
// Deterministic cache parameters for each level
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_001Dh(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_001Dh(uint32_t Leaf) const {
// This is nearly a copy of CPUID function 4h
// There are some minor changes though
@@ -1201,12 +1180,12 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_001Dh(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_Reserved(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_Reserved(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
return Res;
}
FEXCore::CPUID::XCRResults CPUIDEmu::XCRFunction_0h() {
FEXCore::CPUID::XCRResults CPUIDEmu::XCRFunction_0h() const {
// This just returns XCR0
FEXCore::CPUID::XCRResults Res{
.eax = static_cast<uint32_t>(XCR0),
Loaded 100 of 634 files, more files were not shown because too many files have changed in this diff. Show more