Compare commits

..
585 Commits
Author SHA1 Message Date
Ryan Houdek 4a7839b5ac Docs: Update for release FEX-2311.1 2023-11-11 11:59:57 -08:00
Ryan Houdek d8efcb39b8 FEX: Only pass CPU tunables to FEXCore and FEXLoader
This fixes an issue where CPU tunables were ending up in the thunk
generator which means if your CPU doesn't support all the features on
the *Builder* then it would crash with SIGILL. This was happening with
Canonical's runners because they typically only support ARMv8.2 but we
are compiling packages to run on ARMv8.4 devices.

cc: FEX-2311.1
2023-11-11 11:58:25 -08:00
Ryan Houdek fa8c35feba Docs: Update for release FEX-2311 2023-11-07 09:42:16 -08:00
Alyssa Rosenzweig 65b7d4007e Merge pull request #3256 from Sonicadvance1/unittest_for_3254
unittests/ASM: Adds unittest for implicit flag clobber for #3254
2023-11-07 10:02:07 -04:00
Ryan Houdek 2c0444e846 unittests/ASM: Adds unittest for implicit flag clobber for #3254
It looks like currently FEX has a bug around implicit flag clobbering
with this pull request where an IR operation that implicitly clobbers
flags isn't correctly saving the NZCV flags before doing the operation.

Adds a unit test that specifically captures this issue. RAX will be 1 or
0 depending on if the flags are clobbered incorrectly or not.
2023-11-07 02:04:11 -08:00
Ryan Houdek b92e716d0c Merge pull request #3255 from Sonicadvance1/unittest_flags_signals
FEXLinuxTests: Adds a unittest for eflags and signals around a inlined syscall
2023-11-06 05:43:44 -08:00
Ryan Houdek a7a1365cf7 FEXLinuxTests: Adds a unittest for eflags and signals around a inlined syscall
When attempting to debug #3162 I had noticed spurious behaviour around
what I assumed to be eflags getting corrupt around inlined syscalls.
This turned out to be a red herring but to ensure we are still testing
this, create a fully fleshed out unit test.

This test ensures a couple of things.
1) A flag that is set or unset before a syscall doesn't have its data
   corrupt
2) An inline syscall doesn't corrupt the eflags, checking the eflag
   result after returning from the syscall.
3) A signal occuring while in an inline syscall returns the correct
   eflags information in the signal handler information

This test gets accomplished by setting or unsetting a particular flag
and then calling the futex syscall in a way that is guaranteed to be
inlined and also wait forever. Then the parent thread will signal with a
SIGTERM and read back the signal information. It does this multiple
times for each flag we care about.
2023-11-06 05:33:03 -08:00
Alyssa Rosenzweig 8ee5b5cf50 Merge pull request #3253 from Sonicadvance1/fix_double_munmap
JITArm64: Fixes double munmap issue that was causing crashes
2023-11-06 07:57:01 -04:00
Alyssa Rosenzweig b45023bedf Merge pull request #3251 from Sonicadvance1/fexlinuxtests_use_intel
FEXLinuxTests: Compile tests with masm=intel
2023-11-06 07:56:29 -04:00
Alyssa Rosenzweig 3702e513f5 Merge pull request #3252 from Sonicadvance1/fix_fillstaticregisters
ARMEmitter: Fix GPR fill mask in `FillStaticRegs`
2023-11-06 07:50:58 -04:00
Ryan Houdek 829384e488 JITArm64: Fixes double munmap issue that was causing crashes
While tracking issues in #3162, I had encountered a random crash that I
started hunting. It was very quickly apparent that this crash was
unrelated to that PR. I just happened to be running a unittest that was
creating and tearing down a bunch of threads that exacerbated the
problem.

See as follows with the strace output:
```
[pid 269497] munmap(0x7fffde1ff000, 16777216) = 0
[pid 269497] munmap(0x7fffde1ff000, 16777216 <unfinished ...>
[pid 268982] mmap(NULL, 16777216, PROT_READ|PROT_WRITE|PROT_EXEC, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7fffde1ff000
[pid 269497] <... munmap resumed>)      = 0
```

One thread is freeing some memory with munmap, another one then does a mmap and gets the same address back.
Nothing too crazy at initial glance, but taking a closelier look, we can
see that there are two strange oddities:
1) We are double unmapping the same address range through munmap
2) The second munmap is interrupted and returns AFTER the mmap.

This has the unfortunate side-effect that the mmap that just returned
the same address has actually just been unmapped! This was resulting in
spurious crashes around thread creation that was SUPER hard to nail
down.

The problem comes down to how code buffer objects are managed, in
particular how the Arm64Emitter and Dispatcher handled its buffers.

Arm64Emitter is inherited by two classes; Dispatcher, and Arm64JITCore.
On class destruction the emitter would free its internal tracking
buffer. Additionally on destruction, the Arm64JITCore would walk through
all of its CodeBuffers and free them. The problem ends up being that in
the Arm64JITCore, it would free its code buffers which also ended up
being the current active buffer bound to the Arm64Emitter. Thus causing
the Arm64Emitter to come back around and try to free the same buffer
again.

This is a double-free problem! and was only visible on thread exiting!
Can't track double frees with mmap and munmap with current tooling!

This problem typically didn't occur because of how fast the destruction
usually takes and jemalloc inbetween also typically means the problem
doesn't occur. Initially thinking this was a threaded pool allocator bug
because typically the new allocation would end up in there once a new
thread was spinning up.

Now we change behaviour, Arm64Emitter doesn't do any buffer management
itself, instead just passing an initial buffer on to its internal buffer
tracking if given one up front.

This leaves the Dispatcher and the Arm64JITCore to do their buffer
management and ensuring there is no double free.

The day is saved!
2023-11-05 19:18:40 -08:00
Ryan Houdek ed23fbe932 InstCountCI: Update for SRA spill changes. 2023-11-05 02:59:19 -08:00
Ryan Houdek 0af0427efd ARMEmitter: Fix GPR fill mask in FillStaticRegs
This mask was being used incorrectly, it's a GPR spill mask for host
GPRs not an index in to the SRA array. Search the array of SRA registers
for the first one in the mask first to use as a temporary.

Fixes an issue with 32-bit inline syscalls where the first register
being spilled was r8, which was beyond the size of SRA registers on
32-bit processes. This would cause FEX to read the value just after
x32::SRA which is x32::RA. This would mean it would use r20 as a
temporary, corrupting the register in the process.

I noticed this while poking at #3162, but also when I was looking at a
memory buffer ownership problem.
2023-11-05 02:54:48 -08:00
Ryan Houdek a499272d81 FEXLinuxTests: Compile tests with masm=intel
ATT is maddening and a unittest I'm writing was going to be a nightmare.
2023-11-04 22:00:45 -07:00
Ryan Houdek e91c5ff906 Merge pull request #3249 from Sonicadvance1/gdbserver_frontend_prep
GDBServer: Preparation work to get this moved to the frontend
2023-11-03 20:21:28 -07:00
Ryan Houdek 190f7c27e0 Merge pull request #3243 from Sonicadvance1/loadfile_unsized
FEXCore/FileLoading: Updates helper to load file that is backed by memory
2023-11-03 20:21:12 -07:00
Ryan Houdek b15f0b5d36 FEXCore/FileLoading: Updates helper to load file that is backed by memory
When attempting to read files that aren't backed by a filesystem then
our current read file helpers fail since they query the file size
upfront.

Change the helper so that it doesn't query the size and just reads the file if it
can be opened. This lets us read `/proc/self/maps` using helpers.
2023-11-03 07:01:39 -07:00
Ryan Houdek e2c65189ff GDBServer: Preparation work to get this moved to the frontend
GDBServer is inherently OS specific which is why all this code is
removed when compiling for mingw/win32. This should get moved to the
frontend before we start landing more work to clean this interface up.

Not really any functional change.

Changes:

FEXCore/Context: Adds new public interfaces, these were previously
private.
- WaitForIdle
   - If `Pause` was called or the process is shutting down then this
     will wait until all threads have paused or exited.
- WaitForThreadsToRun
   - If `Pause` was previously called and then `Run` was called to get
     them running again, this waits until all the threads have come out
     of idle to avoid races.
- GetThreads
   - Returns the `InternalThreadData` for all the current threads.
   - GDBServer needs to know all the internal thread data state when the
     threads are paused which is what this gives it.

GDBServer:
- Removes usages of internal data structures where possible.
   - This gets it clean enough that moving it out of FEXCore is now
     possible.
2023-11-02 20:11:01 -07:00
Ryan Houdek 03f63f99a8 FEXCore: Moves StringUtils to FEXCore headers
Once gdbserver gets moved to the frontend this will need to be in the
includes.
2023-11-02 20:09:12 -07:00
Ryan Houdek e305a9a0d5 Merge pull request #3248 from neobrain/fix_thunks_async_callback_error
Thunks: Print error if guest-provided callbacks are called asynchronously
2023-11-02 12:03:22 -07:00
Tony Wasserka 3a90dbbb35 Thunks: Print error if guest-provided callbacks are called asynchronously from the host 2023-11-02 19:50:58 +01:00
Ryan Houdek 5103f2d92b Merge pull request #3247 from alyssarosenzweig/refactor/nzcv-prereq
Preparatory patches for nzcv
2023-11-01 14:07:14 -07:00
Alyssa Rosenzweig d4a6b031ea Merge pull request #3245 from Sonicadvance1/remove_gdbpausecheck
FEXCore: Removes gdb pause check handler
2023-11-01 15:46:40 -04:00
Alyssa Rosenzweig 319cf4bf3d Arm64: Preserve flags in ExitFunction
Mostly harmless (except for fusion on cortexes), prepares us for nzcv work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Alyssa Rosenzweig d75c0f2c50 IR: Add missing flag clobbers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Alyssa Rosenzweig a586d3823d IR: VFCMPEQ does not clobber flag
Oversight.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Alyssa Rosenzweig 5522c6db9c OpcodeDispatcher: Use jump wrappers
Mostly automated replacement + renaming for build fixing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Alyssa Rosenzweig 367e1658ad OpcodeDispatcher: Add jump wrappers
These should always be used in the dispatcher rather than the raw jumps they
translate to, as they ensure that flags are flushed. Eliminates a class of bugs
that will become a lot easier to hit with the new nzcv work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Alyssa Rosenzweig bbad06f81a ArchHelpers: Add cfinv()
FlagM goodness.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:35 -04:00
Mai 9612b2fe4b Merge pull request #3240 from Sonicadvance1/optimize_palignr_zero
OpcodeDispatcher: Optimize palignr with zero immediate
2023-11-01 05:55:52 +01:00
Mai ef5503f0b7 Merge pull request #3239 from Sonicadvance1/optimize_blendw
OpcodeDispatcher: Optimize pblendw
2023-11-01 05:54:31 +01:00
Ryan Houdek 26bf67ca76 InstCountCI: Update for pblendw 2023-10-31 21:20:05 -07:00
Ryan Houdek e5df636efd OpcodeDispatcher: Optimize pblendw
Requires #3238 to be merged first since this uses the tbx IR operation.

Worst case is now a three instruction sequence of ldr+ldr+tbx.
Some operations are special-cased, which definitely doesn't cover all
possible cases we could use without tbx, but as a worst case improvement
this is a significant improvement.
2023-10-31 21:18:02 -07:00
Ryan Houdek f34f4a0227 unittests/ASM: Extend pblendw test for all swizzles
We can't sanely test all 256 swizzle masks, so walk the array of random
data for each swizzle and CRC the results each step along the way.
This will give us a final result rather than each individual step.
2023-10-31 21:16:19 -07:00
Mai f45722d2cd Merge pull request #3238 from Sonicadvance1/optimize_blendps
OpcodeDispatcher: Optimize blendps
2023-11-01 04:32:33 +01:00
Ryan Houdek db7ef0e4b0 InstCountCI: Update for blendps
All sixteen swizzles are now optimal.
2023-10-31 20:14:56 -07:00
Ryan Houdek de10cbad98 OpcodeDispatcher: Optimize blendps
A bunch of blendps swizzles weren't optimal. This optimizes all swizzles
to be optimal.

Two instructions can be more optimal without a tbx but the rest required
tbx to be optimal since they don't match ARM's swizzle mechanics.
2023-10-31 20:06:45 -07:00
Ryan Houdek e4d9c264d8 IR: Adds support for tbx 2023-10-31 20:06:45 -07:00
Ryan Houdek 97b3efa90a unittests/ASM: Update blendps to test all swizzles
There are sixteen swizzles so it can fit in a single test.
2023-10-31 20:06:45 -07:00
Mai 18065199e3 Merge pull request #3242 from Sonicadvance1/gdbserver_threadnames
GdbServer: Fixes returning thread names
2023-11-01 04:02:14 +01:00
Mai 77d92872bc Merge pull request #3212 from Sonicadvance1/dpp_opt
OpcodeDispatcher: Optimize 128-bit DPPS and DPPD
2023-11-01 04:01:05 +01:00
Ryan Houdek 460f13be71 FEXCore: Removes gdb pause check handler
gdbserver is currently entirely broken so this doesn't change behaviour.
The gdb pause check that we originally had an excessive amount of
overhead.
Instead use the pending interrupt fault check that was wired
up for wine.
This makes the check very lightweight and makes it more reasonable to
implement a way to have gdbserver support attaching to a process.
2023-10-31 18:40:00 -07:00
Ryan Houdek 47c9463217 GdbServer: Fixes returning thread names
gdb gets angry if we return text with `<No Name>` in an xml text field.
Instead only return a name if we have one and gdb will take care of the
rest.

Additionally change the formatting of the return packet, it doesn't need
the xml version header.
2023-10-29 17:16:46 -07:00
Ryan Houdek 15c825f362 Merge pull request #3241 from neobrain/refactor_thunks_drop_async
Thunks/xcb: Drop unused and incomplete support for asynchronous callbacks
2023-10-29 13:27:29 -07:00
Alyssa Rosenzweig bbd20b47ba Merge pull request #3236 from Sonicadvance1/fix_strenum
Config: Fixes string enum parser with multiple arguments
2023-10-29 07:32:57 -04:00
Alyssa Rosenzweig 5b70209728 Merge pull request #3233 from Sonicadvance1/print_vixl
JIT: Implements Print support for vixl sim
2023-10-29 07:32:32 -04:00
Alyssa Rosenzweig a287f2a189 Merge pull request #3237 from Sonicadvance1/round_classification
InstCountCI: Update rounds{s,d} classification
2023-10-29 07:31:22 -04:00
Tony Wasserka 3a240e3b61 Thunks/xcb: Drop unused and incomplete support for asynchronous callbacks 2023-10-29 11:05:24 +01:00
Ryan Houdek de1e593ec2 InstCountCI: Update for palignr zero immediate 2023-10-27 15:11:07 -07:00
Ryan Houdek 13cd8b33a2 OpcodeDispatcher: Optimize palignr with zero immediate
These turns in to moves
2023-10-27 15:09:31 -07:00
Ryan Houdek 0f26bc20a3 unittests/ASM: Adds palignr tests for zero immediate
These effectively turn in to moves.
2023-10-27 15:07:52 -07:00
Ryan Houdek eacab3cc22 InstCountCI: Update rounds{s,d} classification
This is optimal without AFP. We already test these with AFP in a
different file.
2023-10-27 12:38:02 -07:00
Ryan Houdek 09e3371a0d Config: Fixes string enum parser with multiple arguments
Messed up when originally implementing this, substr's second argument is
requested substring length, not the ending position.

Noticed this while trying to parse multiple FEX_HOSTFEATURES options.
2023-10-27 12:04:04 -07:00
Alyssa Rosenzweig ff3f7345b6 Merge pull request #3235 from Sonicadvance1/cpu_names
CPUID: Adds some missing cpu core names
2023-10-27 08:00:43 -04:00
Alyssa Rosenzweig 8181e53727 Merge pull request #3234 from Sonicadvance1/minor_bfxil
Arm64: Minor optimization to bfxil and bfi
2023-10-27 08:00:28 -04:00
Alyssa Rosenzweig 5431aa5a28 Merge pull request #3232 from Sonicadvance1/assert_code
IR: Print assert code for IR EmitValidation
2023-10-27 07:58:51 -04:00
Ryan Houdek 1a293cc542 CPUID: Adds some missing cpu core names
EZ-PZ
2023-10-26 21:03:32 -07:00
Ryan Houdek b1e78934ad InstCountCI: Improvement for bfi/bfxil 2023-10-26 20:26:48 -07:00
Ryan Houdek 61f22911c7 Arm64: Minor optimization to bfxil and bfi
When the destination doesn't alias the source, we can remove a final mov
from both of these operations.

Does some minor code improvement.
2023-10-26 20:25:44 -07:00
Ryan Houdek fe8778bb96 JIT: Implements Print support for vixl sim
Everytime I want to quickly output a value for testing I tend to use
Print which didn't work under the simulator.
Give this a quick fix to wire up the jump to the vixl sim.
2023-10-26 18:34:47 -07:00
Ryan Houdek 14e5ea1e22 Merge pull request #3230 from Sonicadvance1/optimize_atomic_fetch
IR: Optimize unused result atomic fetch mop to just atomic mop
2023-10-26 15:57:20 -07:00
Ryan Houdek 74f1205f33 IR: Print assert code for IR EmitValidation
Currently we don't get why an IR emit failed in the assert message. Put
the code in to the message so it is easier to see.
This also resolved the issue that when in RelWithDebInfo the assert line
would typically be the end of the IR emission function, so you couldn't
see which assert actually triggered. Now since the message is printed
this is easier

Before:
```
[ASSERT]
```

After:
```
[ASSERT] Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit
```

This is a lot easier and better data than what #3227 proposed.
2023-10-26 15:51:56 -07:00
Ryan Houdek 4045bfd187 Merge pull request #3229 from Sonicadvance1/instcountci_3dnow
OpcodeDispatcher: Optimize a few 3DNow! operations
2023-10-26 15:42:50 -07:00
Ryan Houdek 9db93a43dd IR: Adds comments that atomic op layout must match
If the fetch and nonfetch versions mismatched then the DCE optimization
which changes the IR operation would break handily.

Add a comment as a reminder so if anyone touches this they will
understand.
2023-10-26 15:40:55 -07:00
Ryan Houdek a379d50729 Merge pull request #3226 from neobrain/fix_thunks_incomplete_but_compatible
Thunks: Skip data layout analysis for types that are always assumed compatible
2023-10-26 13:44:31 -07:00
Alyssa Rosenzweig f5822f83b0 Merge pull request #3231 from Sonicadvance1/alu_ops_more_garbage
OpcodeDispatcher: Allow garbage in upper bits for more ALU ops
2023-10-26 06:23:14 -04:00
Ryan Houdek 5028434292 IR: Optimize unused result atomic fetch mop to just atomic mop
When the result of an atomic fetch operation is unused then we can
safely convert it to a non-fetching version of the operation.

This happens hundreds of times per process as far as I can tell. No idea
if this actually helps any hardware, but theoretically it can allow CPUs
to not stall waiting for writeback of the atomic operation.
Couldn't detect any performance improvements in the various little
things I was poking at least. Very trivial to support so add it.

Unaligned variants are already handled in our unaligned fault handler
since the only difference is the acquire semantic is dropped and the
destination register is the zero register.
2023-10-25 21:00:04 -07:00
Ryan Houdek 8f0461cac8 IR: Adds support for non-fetch Atomci CLR and NEG 2023-10-25 21:00:04 -07:00
Ryan Houdek 0b4cb23411 InstCountCI: Update for more garbage data 2023-10-25 20:46:31 -07:00
Ryan Houdek bab96b9441 OpcodeDispatcher: Allow garbage in upper bits for more ALU ops
Secondary ALU operations were missed and when the operation is 4-bytes
in size we can also allow garbage upper bits since the JIT will emit a
32-bit operation for this instruction which is safe.

Optimizes some bad codegen around 32-bit ALU operations.
2023-10-25 20:44:36 -07:00
Ryan Houdek 11e2f14185 InstCountCI: Update for 3DNow! changes 2023-10-25 19:30:13 -07:00
Ryan Houdek 6177290e9d OpcodeDispatcher: Optimize a few 3DNow! operations
Instead of using VInsElement in pi2fw and pf2iw, just use uzp1 to ensure
we don't unintentionally add to RA pressure.

Additionally we can generate the constant needed for pmulhrw directly
using the movi instruction. Converts two instructions in to one.

Under FEX's current constraints this makes all 3DNow! instructions
optimal.
2023-10-25 19:27:54 -07:00
Ryan Houdek 20a54913bd IR: Support shifted imms in VectorImm
We will want to use this in some cases.
2023-10-25 19:27:00 -07:00
Tony Wasserka dd9ed89a7a Thunks/X11: Drop unneeded type annotation 2023-10-25 20:38:57 +02:00
Tony Wasserka a4e1e0a1fb Thunks/gen: Skip data layout analysis for types that are always assumed compatible 2023-10-25 19:35:02 +02:00
Tony Wasserka 5e9f69001d Thunks/gen: Clarify pointer parameter handling
One of the subconditions was always true, so it can safely be removed.
2023-10-25 19:35:02 +02:00
Ryan Houdek 39c5ab1c81 Merge pull request #3225 from neobrain/fix_32bit_funcptrs
Thunks: Fix function pointer support on 32-bit
2023-10-25 07:07:18 -07:00
Ryan Houdek 5bf790324c Merge pull request #3224 from neobrain/feature_thunk_strict_everywhere
Thunks: Annotate pointer parameters throughout all thunked libraries
2023-10-25 07:00:42 -07:00
Tony Wasserka adcdb32d49 Thunks: Fix function pointer support on 32-bit 2023-10-25 14:37:22 +02:00
Tony Wasserka f264578f12 Thunks: Unconditionally enable strict processing mode 2023-10-25 12:39:57 +02:00
Tony Wasserka 7149da387a Thunks: Annotate pointer parameters throughout all thunked libraries 2023-10-25 12:39:57 +02:00
Alyssa Rosenzweig 0f3d14e7c0 Merge pull request #3223 from Sonicadvance1/testharness_page_size
TestHarnessRunner: Don't hardcode stack allocation to 4096 bytes
2023-10-24 10:52:35 -04:00
Ryan Houdek aad5080224 TestHarnessRunner: Don't hardcode stack allocation to 4096 bytes
Just allocate a single page that we query at runtime.
2023-10-24 07:36:13 -07:00
Ryan Houdek c77a3d673c Merge pull request #3222 from Sonicadvance1/nzcv_opt_bug
unittests: Adds test for bug from #3162
2023-10-23 18:19:16 -07:00
Ryan Houdek 0ff2e6e1e3 unittests: Adds test for bug from #3162
This PR has a bug around flags calculation and REP LODS{B,W,D,Q}.
This currently passes on main but fails on #3162.

Bug only occurs in 32-bit instead of 64-bit with the same test. Should
help diagnose the bugs in #3162.
2023-10-23 16:55:56 -07:00
Ryan Houdek 8538f5bac4 Merge pull request #3221 from Sonicadvance1/instcountci_missing_insts
InstCountCI: Adds two missing variants of movd/movq
2023-10-23 15:26:58 -07:00
Ryan Houdek a305baf6e5 InstCountCI: Fixes some mislabeled instructions
These are optimal according to our standards.
2023-10-23 15:15:03 -07:00
Ryan Houdek 807619aa02 InstCountCI: Adds two missing variants of movd/movq
We only had the move to memory destination version, ensure we test to
GPR as well.
2023-10-23 15:01:18 -07:00
Ryan Houdek 6db2125b41 Merge pull request #3220 from Sonicadvance1/override_flagm
InstCountCI: Support disabling flagm extensions
2023-10-23 14:45:45 -07:00
Ryan Houdek 9f6d80fe5d InstCountCI: Duplicate tests that change behaviour based on flagm
Necessary for #3162 to have consistent behaviour in CI
2023-10-23 14:03:19 -07:00
Ryan Houdek 423ce12001 InstCountCI: Disable flagm and Flagm2 on tests
Most of these will get duplicated in the next commit
2023-10-23 14:02:50 -07:00
Ryan Houdek 4edd72fc33 InstCountCI: Support disabling flagm extensions
This is necessary so #3162 can give consistent results
2023-10-23 14:02:24 -07:00
Ryan Houdek 978f607dd9 Merge pull request #3177 from neobrain/feature_thunk_pointer_annotations
Thunks: Add new pointer annotations to assist data layout analysis
2023-10-23 13:00:27 -07:00
Ryan Houdek 9ba78c9771 Merge pull request #3219 from Sonicadvance1/enable_vixlsim_instcountci
github: Enables Vixl simulator on x86 host for instcountci
2023-10-23 12:49:33 -07:00
Ryan Houdek e2144345c0 github: Enables Vixl simulator on x86 host for instcountci
Otherwise features get filled out weirdly.
2023-10-23 12:42:40 -07:00
Ryan Houdek 63e4c3682d Merge pull request #3218 from Sonicadvance1/opt_df
OpcodeDispatcher: Optimize DF pointer offset calculation
2023-10-23 11:41:22 -07:00
Ryan Houdek 2956e84ead InstCountCI: Update for constant prop improvement 2023-10-23 10:39:29 -07:00
Ryan Houdek 4466c50c2b ConstProp: Optimize SubShift and Add with negative
When SubShift (LSL) occurs with both sources constant then optimize away
the calculation.

Additionally if add is found to have one immediate constant where the
inverse of the constant fits in to ImmAddSub range, then invert the
constant and change it in to a sub.

This optimizes the cases when direction flag is known upfront in an
instruction.
2023-10-23 10:36:33 -07:00
Ryan Houdek dd5ca1d349 InstCountCI: Update for DF pointer optimization. 2023-10-23 10:14:51 -07:00
Ryan Houdek 95c756b466 OpcodeDispatcher: Optimize DF pointer offset calculation
Previously this moved two constant, did a compare and a csel. Four
instructions in total. It also corrupts NZCV which we want to use for
other things.

This new codegen emits one constant and one subtract instruction, two
instructions total and doesn't touch NZCV.

More optimal!
2023-10-23 09:27:41 -07:00
Ryan Houdek 99465faf63 IR: Implements support for subtract with shifted register
Will be used soon.
2023-10-23 09:27:41 -07:00
Ryan Houdek e018917f76 Merge pull request #3216 from alyssarosenzweig/opt/nzcv-infra
Prep commits for NZCV modelling
2023-10-23 07:39:47 -07:00
Ryan Houdek a261d9909e Merge pull request #3214 from Sonicadvance1/fix_bug
FEXCore: Fixes bug in vector `ZextAndMaskingElimination` pass
2023-10-23 07:29:43 -07:00
Alyssa Rosenzweig 6de8bc6848 IR: Annotate instructions with implicit flag clobber
Audit the code base and mark any instruction that implicitly clobbers flags so
it can get special handling in the dispatcher to spill NZCV ahead of emitting.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:21:47 -04:00
Alyssa Rosenzweig d87155e4ee IR: Add infrastructure for modelling flag clobbers
Lots of instructions clobber NZCV inadvertently but are not intended to write to
the host flags from the IR point-of-view. As an example, Abs logically has no
side effects but physically clobbers NZCV due to its cmp/csneg impl on non-CSSC
hw. Add infrastructure to model this in the IR so we can deal with it when we
start using NZCV for things.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:21:47 -04:00
Alyssa Rosenzweig 7484cacaf9 InstCountCI: Update for VInsertElement change
Only SVE256 codepath affected.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:21:47 -04:00
Alyssa Rosenzweig 42259974c4 Arm64: Preserve NZCV in VInsertElement
So we don't need to mark VInsertElement as implicit clobber in the common case.
Only afects sve256 which doesn't exist yet.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:21:34 -04:00
Alyssa Rosenzweig e455996dbd OpcodeDispatcher: Remove silly shift branching
The flag generation code does this internally and more efficiently.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:16:38 -04:00
Alyssa Rosenzweig b5dd1d05e9 Dispatcher: Preserve NZCV
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:16:38 -04:00
Alyssa Rosenzweig bbaf70da15 Dispatcher: Yeet pointless subs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:16:29 -04:00
Ryan Houdek 65e8d094ef Merge pull request #3215 from alyssarosenzweig/hack/16k-units
FEXLoader: Query runtime page size
2023-10-23 06:43:43 -07:00
Alyssa Rosenzweig 4c801d594a FEXLoader: Query runtime page size
This lets most of the ASM tests run on 16K Linux hosts which is good because I
have a Mac and I'm bad at computer.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 09:35:22 -04:00
Ryan Houdek d4403edea9 OpcodeDispatcher: Updates COMIS to eliminate scalar moves
This was one of the few things that managed to hit the previously
removed optimization. Just fix the OpcodeDispatcher instead.
2023-10-21 21:33:07 -07:00
Ryan Houdek 06ef012fb2 FEXCore: Fixes bug in vector ZextAndMaskingElimination pass
With the previous RCLSE pass optimization that fixes store->load
forwarding, this pass started optimizing harder.

This hit a bug with this vmov removal that previously didn't get hit.
In particular this would eliminate vmov IR operations even if they were
zero extending a vector.
Since we have dramatically cleaned up the amount of vmov IR operations
we are generating, remove this optimization entirely. In the games I
tested, the only game that hit this "optimization" was Ender Lilies and
it started generating broken code for the single block of instructions
that did.

Adds a unit test for this case just in-case it comes back in the future
for some reason.

Fixes an issue where Ender Lilies would flash the screen to black every
time an enemy hit the player character.
2023-10-21 21:21:14 -07:00
Mai 8f8f37684a Merge pull request #3213 from Sonicadvance1/fix_repres
JITArm64: Fixes bug in rpres scalar operations
2023-10-22 05:05:14 +02:00
Ryan Houdek 7140b8d901 InstCountCI: Update for RPRES fix 2023-10-21 15:29:11 -07:00
Ryan Houdek d5beba9423 JITArm64: Fixes bug in rpres scalar operations
Noticed this during code investigation, these two operations were
swapped.

Would have caused issues if anything supported RPRES today.
2023-10-21 15:24:43 -07:00
Ryan Houdek 826e15aea9 unittests/ASM: Adds dpps/dppd broadcast mask tests
Ensures that the optimization around the broadcast mask is correct.
2023-10-20 18:15:43 +02:00
Ryan Houdek 14e80ce228 InstCountCI: Update for DPPS/DPPD
Adds some new destination broadcast masks to ensure we handle most of
them.
2023-10-20 18:15:43 +02:00
Ryan Houdek 165d3d3d4d Arm64JIT: Fixes VDupElement so it respects 64-bit vector duping
In some cases when we want the upper bits to be zero, this is the
desired behaviour
2023-10-20 18:15:43 +02:00
Ryan Houdek 887200e571 OpcodeDispatcher: Optimize 128-bit DPPS and DPPD
These instructions aren't super amazing due to the fact that they have
both a source mask and a destination duplication mask.

Setup a case where we can generate more optimal code in /most/ cases.

There are a few that still fall down a "bad" path for the result
broadcast but in most cases they are optimal. Still to be seen what
games typically use the broadcast mask as.

AVX in its infinite wisdom expanded DPPS to 256-bit, while leaving DPPD
to only support 128-bit still. This leaves the original implementation
alone for 256-bit DPPS since I don't want to break it.

This is another instruction that gets a free optimization when
SVE-128bit is supported!
2023-10-20 18:02:27 +02:00
Ryan Houdek 2c0bc0654d IR: Adds new VFAddV operation
SVE added this instruction natively, we can take advantage of it on
SVE-128bit systems which is quite nice.

Will be used soon.
2023-10-19 16:38:11 +02:00
Ryan Houdek b3d76bd2f1 IR: Adds DPPS and DPPD source masks
This will get used for these instructions soon
2023-10-19 16:36:19 +02:00
Ryan Houdek 2e694412f4 Merge pull request #3211 from lioncash/ext
VectorOps: Handle SVE VExtr a little better
2023-10-19 16:04:02 +02:00
Lioncache d84577c36c VectorOps: Handle SVE VExtr a little better
If the source registers don't alias the destination, then we can
safely move the lower bits over to it without using a temporary.
2023-10-19 15:11:23 +02:00
Ryan Houdek cf9c2aa72c Merge pull request #3206 from Sonicadvance1/fix_syscall
Linux: Fixes issue with *at syscalls with absolute paths not working
2023-10-19 15:05:34 +02:00
Ryan Houdek 1cb8e4891c Merge pull request #3210 from lioncash/fcadd
VectorOps: Handle SVE VFCADD a little better
2023-10-19 15:05:15 +02:00
Lioncache 24f2796141 VectorOps: Handle SVE VFCADD a little better
If no registers alias, then we can move the first source directly into the
destination and then perform the FCADD operation as opposed to using a
temporary.
2023-10-19 14:48:46 +02:00
Tony Wasserka cb215b5f21 FEXLinuxTests/thunks: Add assume_compatible_data_layout tests 2023-10-19 12:49:00 +02:00
Tony Wasserka 0cf2695772 FEXLinuxTests/thunks: Add tests for opaque types 2023-10-19 12:49:00 +02:00
Tony Wasserka 6a6886305e Thunks/gen: Add assume_compatible/is_opaque annotations
These annotations allow for a given type or parameter to be treated as
"compatible" even if data layout analysis can't infer this automatically.

assume_compatible_data_layout is more powerful than is_opaque, since it
allows for structs containing members of a certain type to be automatically
inferred as "compatible".

Conversely however, is_opaque enforces that the underlying data is never
accessed directly, since non-pointer uses of the type would still be
detected as "incompatible".
2023-10-19 12:49:00 +02:00
Tony Wasserka 5ef7537e61 unittests/thunks: Add ptr_passthrough tests 2023-10-19 12:49:00 +02:00
Tony Wasserka 167fe85cc3 Thunks: Implement ptr_passthrough annotation
This annotation can be used for data types that can't be repacked
automatically even with custom repack annotations. With ptr_passthrough,
the types are wrapped in guest_layout and passed to the host like that.
2023-10-19 12:49:00 +02:00
Tony Wasserka cf65747667 Thunks: Introduce an intermediate guest_layout wrapper to unpack callback arguments
This will be used later to aid automatic struct repacking.
2023-10-19 12:48:59 +02:00
Tony Wasserka 27bb28b47f Thunks: Carry annotations in callback wrappers of host functions
Previously, two functions with the same signature would always be wrapped
in the same logic. This change allows customizing one function with
annotations while leaving the other one unchanged.
2023-10-19 12:48:59 +02:00
Tony Wasserka a00da800e7 Thunks: Rename funcptr_types to thunked_funcptrs
This reflects its purpose slightly better, particularly since future patches
will add more information to this object.
2023-10-19 12:48:59 +02:00
Tony Wasserka bf835e80ac Thunks: Bump compiler requirements to C++20 2023-10-19 12:48:59 +02:00
Tony Wasserka 8f246b206b Merge pull request #3209 from neobrain/refactor_revert_vulkan_reorder 2023-10-19 12:45:14 +02:00
Ryan Houdek 3c5c23bf36 Merge pull request #3208 from lioncash/avg
VectorOps: Handle SVE VURAvg a little better
2023-10-19 12:38:56 +02:00
Tony Wasserka 5bcfaf4b9f Thunks/vulkan: Revert reordering changes from 180d16af7a
These interfere heavily with ongoing work. Let's reapply the reordering
once the dust has settled instead.
2023-10-19 12:31:33 +02:00
Lioncache 1f6c6345d9 VectorOps: Handle SVE VURAvg a little better
We can perform less moves by checking for scenarios where aliasing
occurs. Since addition is commutative (usually, general-case anyway),
order of inputs doesn't strictly matter here.
2023-10-19 12:14:12 +02:00
Ryan Houdek 93792577eb Merge pull request #3207 from lioncash/div
VectorOps: Handle SVE VFDiv a little better
2023-10-19 11:53:45 +02:00
Lioncache 3d23cd5765 VectorOps: Handle SVE VFDiv a little better
In the event no source vectors alias the destination,
we can just move the first source vector into it and
then perform the divide without needing to move afterword.
2023-10-19 11:45:35 +02:00
Ryan Houdek fcc239552c Linux: Fixes issue with *at syscalls with absolute paths not working
When a syscall from the *at series is provided an FD but the path is
absolute then dirfd should be ignored. We weren't correctly doing this.
Now if the path is absolute, but set the argument to the special
AT_FDCWD..
Fixes #3204
2023-10-19 09:48:50 +02:00
Ryan Houdek 8238de024f Merge pull request #3205 from lioncash/max
VectorOps: Handle SVE VSMax/VSMin and VUMax/VUMin paths a little better
2023-10-18 19:24:35 +02:00
Lioncache 39e658f02a VectorOps: Handle more VUMin SVE cases better
We can avoid needing to use movprfx here by moving
directly into the destination when possible and just
doing the UMIN directly
2023-10-18 18:48:13 +02:00
Lioncache e89dd27f2a VectorOps: Handle more VSMin SVE cases better
We can avoid needing to use movprfx here by moving
directly into the destination when possible and just
doing the SMIN directly.
2023-10-18 18:48:13 +02:00
Lioncache f85fae0041 VectorOps: Handle more VUMax SVE cases better
We can avoid needing to use movprfx here by moving
directly into the destination when possible and just
doing the UMAX directly.

Also expands the unsigned max tests to test values with
the sign bit set to ensure all behavior is caught.
2023-10-18 18:48:12 +02:00
Lioncache 65eec673fc VectorOps: Handle more VSMax SVE cases better
Since SMAX performs a comparison and returns the max value regardless
of how the operands are provided, we can check for when the second
input aliases the destination.
2023-10-18 18:48:03 +02:00
Ryan Houdek 5c93a085d2 Merge pull request #3203 from lioncash/movs
OpcodeDispatcher: Handle SSE vector moves into themselves a little better
2023-10-18 16:28:45 +02:00
Lioncache 4b356a7c2c OpcodeDispatcher: Have MOVNTSD go down the non-temporal path
For some reason this was using the regular unaligned path.
2023-10-18 14:59:02 +02:00
Lioncache 2b67f87054 OpcodeDispatcher: Handle SSE vector moves into themselves a little better
Obviously, it's silly to do this, but we should still be generating
optimal code for this case (which is none at all).
2023-10-18 14:58:57 +02:00
Ryan Houdek 1ea40ae676 Merge pull request #3201 from neobrain/fix_flt_thunks_64bit_only
FEXLinuxTests: Temporarily limit thunk test execution to 64-bit guests
2023-10-18 12:40:29 +02:00
Ryan Houdek e0ef32e0bf Merge pull request #3202 from Sonicadvance1/oopsies_vulkan
Thunks: Oops deleted an entry point
2023-10-18 12:38:05 +02:00
Ryan Houdek a2b53c8eb0 Thunks: Oops deleted an entry point
Moving some entries around I managed to delete one.
Fixes Vulkan thunks.
2023-10-18 12:21:28 +02:00
Tony Wasserka 21b6cccb4e FEXLinuxTests: Temporarily limit thunk test execution to 64-bit guests
Thunking isn't fully functional on 32-bit guests currently, so non-trivial
tests would currently hang in that context.
2023-10-18 12:09:10 +02:00
Tony Wasserka d539829251 FEXLinuxTests: Drop .32/.64 suffixes from test names 2023-10-18 12:09:10 +02:00
Ryan Houdek ef321e4bf8 Merge pull request #3200 from lioncash/mov
OpcodeDispatcher: Remove unnecessary 128-bit truncating moves from StoreResult
2023-10-17 12:12:48 +02:00
Lioncache 47a0f14537 OpcodeDispatcher: Remove unnecessary 128-bit truncating moves from StoreResult
Removes the truncating move that we perform inside the StoreResult
function and instead delegates the responsibility to the instruction
implementations themselves.

This removes a lot of redundant moves that occur on 128-bit variants
of AVX instructions.

Also fixes a weird case where we were handling 128-bit SVE
in VBroadcastFromMem when we already have AdvSIMD instructions
that will perfom the zero-extension behavior for us.
2023-10-17 11:07:04 +02:00
Ryan Houdek 6d39f369b0 Merge pull request #3199 from lioncash/loadops
OpcodeDispatcher: Put extra LoadSource options in a struct
2023-10-16 09:42:27 +02:00
Lioncache 2304cfc530 OpcodeDispatcher: Remove prefixing from MemoryAccessType enum
Since this is an enum class, we don't need to add a prefix.
2023-10-16 03:10:33 +02:00
Lioncache 1a39de4509 OpcodeDispatcher: Put extra LoadSource options in a struct
Allows for easier expansion without needing to expand the function definitons.

Also makes a few usages significantly less verbose and makes specifying
options a little more declarative, rather than having to memorize what
each argument is specifying.
2023-10-15 21:18:00 +02:00
Ryan Houdek efb479f88f Merge pull request #3198 from lioncash/rorx
OpcodeDispatcher: Remove redundant moves from rorx
2023-10-15 17:13:00 +02:00
Lioncache b27bf43901 OpcodeDispatcher: Remove redundant moves from rorx
By allowing junk in the upper bits, we can avoid an unnecessary move,
since we'll be ignoring them in the following ROR instruction anyway.
2023-10-15 17:05:15 +02:00
Ryan Houdek c612fa8f2f Merge pull request #3196 from lioncash/fmt 2023-10-13 10:44:43 -07:00
Lioncache 4ccc40f697 Externals: Update fmt from 10.1.0 to 10.1.1
Notably this bugfix version also introduces support for formatting
std::atomic types and std::atomic_flag.

Also, of course keeps our tracked external up to date.
2023-10-13 13:31:49 -04:00
Ryan Houdek cb53a704ba Merge pull request #3195 from Sonicadvance1/fix_missing_vulkan
Thunks: Fixes missing vulkan definitions
2023-10-11 18:51:52 -07:00
Alyssa Rosenzweig 483423674a Merge pull request #3194 from Sonicadvance1/instcountci_multiinst_tests
InstCountCI: Adds some multi instruction tests
2023-10-11 05:06:55 -04:00
Ryan Houdek 180d16af7a Thunks: Fixes missing vulkan definitions
A couple of games were hitting these. Not sure how they were missed in
PR #3159 but adds the missing one.

Small rearrangement to make this easier as well. Hopefully thunk stuff
lands sooner rather than later to automate this for Vulkan.

Maybe `-isystem` instead of `-I` needs to be used unlike what #2076,
might depend on what is installed on the host system.
2023-10-11 01:49:42 -07:00
Ryan Houdek 1acc038826 InstCountCI: Adds some multi instruction tests
Some simple tests to showcase instructions that we can optimize.
- Back to back pushes could be optimized
- Back to back scalar vector operations can be optimized
- Show with AFP that back to back scalar is already optimal
   - Also ensures we don't break this stage.
2023-10-10 16:45:39 -07:00
Alyssa Rosenzweig cc558fd5dc Merge pull request #3192 from Sonicadvance1/optimize_small_push
OpcodeDispatcher: Optimizes < 32-bit register push
2023-10-10 17:10:34 -04:00
Ryan Houdek 8f04223193 InstCountCI: Adds push changed results
Removes an accidental Test.json file that I use locally.
2023-10-10 12:06:36 -07:00
Ryan Houdek 4be649c44e OpcodeDispatcher: Optimizes < 32-bit register push
We don't need zero in the upper bits for a push.
Makes a couple variants optimal.

Adds missing tests to the 32-bit file, since only 32-bit can push a
32-bit register.
2023-10-10 12:04:41 -07:00
Ryan Houdek 6253f4f708 Merge pull request #3186 from Sonicadvance1/afp_support
IR: Adds scalar vector insert operations
2023-10-10 11:28:22 -07:00
Ryan Houdek cc2eef619c InstCountCI: Update for AFP optimizations
A bunch of random instructions have converted to be optimal in a vacuum.
2023-10-10 03:44:59 -07:00
Ryan Houdek 3bff42e6a7 OpcodeDispatcher: Wire up support for the new scalar insert operations 2023-10-10 03:44:58 -07:00
Ryan Houdek 2671246fef IR: Adds scalar vector insert operations
These IR operations are required to support AFP's NEP mode which does
vector insert in to the destination register. Additionally it gives us
tracking information to allow optimizing out redundant inserts on
devices that don't support AFP natively.

In order to match x86 semantics we need to support binary and unary
scalar operations that do a final insert in to a vector. With optional
zeroing of the top 128-bits for AVX variants.

A tricky thing is that in binary operations this means that the
destination and first source have an intrinsically linked property
depending on if it is SSE or AVX.

SSE example:
- addss xmm0, xmm1
   - xmm0 is both the destination and the first source.
   - This means xmm0[31:0] = xmm0[31:0] + xmm1[31:0]
   - Bits [127:32] are UNMODIFIED.

FEX's JIT jumps through some hoops so that if the destination register
equals the first source register, then it hits the optimal path the
AFP.NEP will insert in to the result. AVX throws a small wrench in to
this due to changed behaviour

AVX example:
- vaddss xmm0, xmm1, xmm2
  - xmm0 is ONLY the destination, xmm1 and xmm2 are the sources
  - This operation copies the bits above the scalar result from the
    first source (xmm1).
  - Additionally this will zero bits above the original 128-bit xmm
    register.
  - xmm0[31:0] = xmm1[31:0] + xmm2[31:0]
  - xmm0[127:32] = xmm1[127:32]
  - ymm0[255:127] = 0

This causes these instructions to support a fairly large table depending
on if the instruction is an SSE or AVX instruction, plus if the host CPU
supports AFP or not.

So while fairly complex, it's handling all the edge cases and gives us
optimization opportunities as we move forward. Currently on non-AFP
supporting devices this has a minor benefit that these IR operations
remove one temporary register, lowering the Register Allocation
overhead.

In the coming weeks I am likely to introduce an optimization pass that
removes redundant inserts because FEX currently does /really/ badly with
scalar code loops.

Needs #3184 merged first.
2023-10-10 03:17:19 -07:00
Ryan Houdek 8cb8f090dd Arm64Emitter: enable/disable AFP on Fill/Spill
When FEX is in the JIT we need to make sure to enable NEP and AH and
then disable when leaving.

Explicitly disabled when the vixl simulator is used since even
attempting to set the bits will cause it to fault out. Ensures
InstCountCI keeps working.
2023-10-10 03:17:18 -07:00
Ryan Houdek a37d89a7d5 HostFeatures: Disable AFP until verified that it is working
Need to audit scalar instruction usage to ensure all uses are okay with
garbage in the upper bits.
2023-10-10 03:17:18 -07:00
Ryan Houdek 252d7712ea Arm64: Save if the host supports AFP 2023-10-10 03:17:18 -07:00
Ryan Houdek c548625fbe InstCountCI: Update tests for disabling AFP
Doesn't change behaviour yet, just prep work.
2023-10-10 03:17:18 -07:00
Ryan Houdek f036a0b84f Merge pull request #3191 from Sonicadvance1/instcountci_multiple
InstCountCI: Support multiple instructions in the tests
2023-10-10 02:53:14 -07:00
Ryan Houdek 2e1389b25e Merge pull request #3184 from Sonicadvance1/armemitter_sized_scalars
ArmEmitter: Adds sized Scalar 1 source and 2 source helpers
2023-10-10 02:53:06 -07:00
Alyssa Rosenzweig a5f82a57fa Merge pull request #3190 from Sonicadvance1/atomic_instcountci
InstCountCI: Adds missing atomic tests
2023-10-10 05:22:36 -04:00
Ryan Houdek cd83d3eb24 InstCountCI: Support multiple instructions in the tests
There are some cases where we want to test multiple instructions where
we can do optimizations that would overwise be hard to see.

eg:
```asm
; Can be optimized to a single stp
push eax
push ebx

; Can remove half of the copy since we know the direction
cld
rep movsb

; Can remove a redundant insert
addss xmm0, xmm1
addss xmm0, xmm2
```

This lets us have arbitrary sized code in instruction count CI, with the
original json key becoming only a label if the instruction array is
provided.

There are still some major limitations to this, instructions that
generate side-effects might have "garbage" after the end of the block
that isn't correctly accounted for. So care must be taken.

Example in the json
```json
"push ax, bx": {
  "ExpectedInstructionCount": 4,
  "Optimal": "No",
  "Comment": "0x50",
  "x86Insts": [
    "push ax",
    "push bx"
  ],
  "ExpectedArm64ASM": [
    "uxth w20, w4",
    "strh w20, [x8, #-2]!",
    "uxth w20, w7",
    "strh w20, [x8, #-2]!"
  ]
}
```
2023-10-09 21:49:53 -07:00
Ryan Houdek 93ab8ab23c InstCountCI: Adds missing atomic tests
This adds all the missing atomic tests in to their own tests files.
This includes all of them except a few choice ones that are in their
original files.

- BTC, BTR, BTS  are in their Secondary/SecondaryGroup files
- CMPXCHG, CMPXCHG8B, CMPXCHG16B are in their Secondary/SecondaryGroup
  files
   - These always imply lock semantics even without the prefix.
2023-10-09 21:18:08 -07:00
Ryan Houdek 462fff2c67 Merge pull request #3189 from Sonicadvance1/remove_warnings_15
FEXCore: Removes a warning about assume discarding side-effects
2023-10-09 17:26:14 -07:00
Alyssa Rosenzweig 8dab35cbf8 Merge pull request #3188 from Sonicadvance1/reconstruct_flags_naming
FEXCore: Renames raw FLAGS location names to signify they can't be used directly
2023-10-09 19:41:09 -04:00
Ryan Houdek a1a479e69f FEXCore: Removes a warning about assume discarding side-effects 2023-10-09 16:04:57 -07:00
Ryan Houdek 6403290019 FEXCore: Renames raw FLAGS location names to signify they can't be used directly
Six of the EFLAGS can't be used directly in a bitmask because they are
either contained in a different flags location or has multiple bits
stored in it.

SF, ZF, CF, OF are stored in ARM's NZCV format in offset 24.
PF calculation is deferred but stored in the regular offset.
AF is also deferred in relation to the PF but stored in the regular
offset.

These /need/ to be reconstructed using the `ReconstructCompactedEFLAGS`
function when wanting to read the EFLAGS.

When setting these flags they /need/ to be set using
`SetFlagsFromCompactedEFLAGS`.

If either of these functions are not used when managing EFLAGs then the
internal representation will get mangled and the state will be
corrupted.

Having a little `_RAW` on these to signify that these aren't just
regular single bit representations like the other flags in EFLAGS should
make us puzzle about this issue before writing more broken code that
tries accessing it directly.
2023-10-08 11:51:11 -07:00
Ryan Houdek 580bd50a00 unittests/ASM: Removes eflags comparison option
This was not used and is also broken.
2023-10-08 11:51:11 -07:00
Ryan Houdek b2a8b0ca12 Merge pull request #3187 from Sonicadvance1/implement_rpres
FEXCore: Implements support for RPRES
2023-10-08 09:57:27 -07:00
Ryan Houdek f78bdf0852 unittests/Emitter: Adds sized scalar unittests. 2023-10-08 09:48:37 -07:00
Ryan Houdek 5652eb4c5d ARMEmitter: Removes templated ptrue/ptrues
Non-templated version exists and templated version gets us nothing.
2023-10-07 23:51:32 -07:00
Ryan Houdek a52bb47551 unittests: Update for rpres optimization 2023-10-07 23:22:51 -07:00
Ryan Houdek 22590dde77 FEXCore: Implements support for RPRES
This allows us to use reciprocal instructions which matches precision of
what x86 expects rather than converting everything to float divides.

Currently no hardware supports this, and even the upcoming X4/A720/A520
won't support it, but it was trivial to implement so wire it up.
2023-10-07 23:13:47 -07:00
Ryan Houdek 6543a80ff9 Merge pull request #3185 from Sonicadvance1/ir_dispatcher_emit
FEXCore/IR: Changes over to automated IR dispatch generation
2023-10-07 21:21:44 -07:00
Ryan Houdek 9c36d1061b Merge pull request #3182 from Sonicadvance1/instcountci_stacking_test_names
InstCountCI: Fixes recursive tests with same filename
2023-10-07 21:21:10 -07:00
Alyssa Rosenzweig 5a3cc7b469 Merge pull request #3183 from Sonicadvance1/instcountci_support_afp_override
InstCountCI: Support overriding AFP features
2023-10-07 19:57:33 -04:00
Ryan Houdek 4cff3e5f1f FEXCore/IR: Changes over to automated IR dispatch generation
Suggested by Alyssa. Adding an IR operation can be a little tedious
since you need to add the definition to JIT.cpp for the dispatch switch,
JITClass.h for the function declared, and then actually defining the
implementation in the correct file.

Instead support the common case where an IR operation just gets
dispatched through to the regular handler. This lets the developer just
put the function definition in to the json and the relevent cpp file and
it just gets picked up.

Some minor things:
- Needs to support dynamic dispatch for {Load,Store}Register and
  {Load,Store}Mem
   - This is just a bool in the json
- It needs to not output JIT dispatch for some IR operations
   - SSE4.2 string instructions and x87 operations
   - These go down the "Unhandled" path
- Needs to support a Dispatcher function override
   - This is just for handling NoOp IR operations that get used for
     other reasons.
- Finally removes VSMul and VUMul, consolidating to VMul
   - Unlike V{U,S}Mull, signed or unsigned doesn't change behaviour here
- Fixed a couple random handler names not matching the IR operation
  name.
2023-10-07 15:01:47 -07:00
Ryan Houdek a1eb571630 ArmEmitter: Adds sized Scalar 1 source and 2 source helpers
Removes the need for an annoying switch statement with scalar operations
for the most part.
2023-10-07 11:51:32 -07:00
Ryan Houdek 559cf6491a InstCountCI: Support overriding AFP features
Also disable AFP under the vixl simulator by default since it doesn't support it.
2023-10-07 11:48:42 -07:00
Ryan Houdek 4bdda1eeb5 InstCountCI: Fixes recursive tests with same filename
This will be used to move AFP tests to a sub-directory
2023-10-07 11:47:16 -07:00
Mai fc70fc3506 Merge pull request #3179 from Sonicadvance1/support_hostfeature_crypto
FEXCore: Support crypto extensions in HostFeatures override
2023-10-06 16:01:59 -04:00
Mai 26ee63cc24 Merge pull request #3181 from Sonicadvance1/remove_spurious_license
External: Remove a spurious license
2023-10-06 15:59:41 -04:00
Ryan Houdek 0092ea7c0b External: Remove a spurious license
This doesn't exist anymore
2023-10-06 09:37:17 -07:00
Ryan Houdek 439a3b9c3a HostFeatures: Use a define 2023-10-06 09:33:41 -07:00
Alyssa Rosenzweig b4ddf36582 Merge pull request #3180 from Sonicadvance1/remove_warnings_14
Linux: Fixes warning in 32-bit clock_settime
2023-10-06 08:13:52 -04:00
Ryan Houdek 12c44f26e5 Linux: Fixes warning in 32-bit clock_settime
This syscall requires a valid pointer otherwise it returns EFAULT.
When going through the glibc helper it can crash before reaching the raw
syscall even.
2023-10-05 17:44:33 -07:00
Ryan Houdek 5b7ba06d5c FEXCore: Support crypto extensions in HostFeatures override
Enables in InstCountCI so Pi users can run InstCountCI can run the tests
without breaking on crypto operations.

When crypto is enabled or disabled just wholesale change AES, CRC32, and
PMULL 128-bit in one step. We don't really care about partial support
here.
2023-10-05 17:41:08 -07:00
Ryan Houdek ee0c1457d8 Docs: Update for release FEX-2310 2023-10-05 14:39:10 -07:00
Alyssa Rosenzweig 3413eb3d98 Merge pull request #3169 from Sonicadvance1/remove_constant_indirection
FEXCore: Support CpuState relative vector named constants
2023-10-05 08:22:48 -04:00
Ryan Houdek 2e0753a244 InstCountCI: Update for named vector constant optimization 2023-10-04 20:57:09 -07:00
Ryan Houdek 8a51bb7a61 FEXCore: Support CpuState relative vector named constants
The motivation towards just having a pointer array in CpuState was that
initialization was fairly cheap and that we have limited space inside
the encoding depending on what we want to do.

Initialization cost is still a concern but doing a memcpy of 128-bytes
isn't that big of a deal.

Limited space in CpuState, while a concern isn't a significant one.
   - Needs to currently be less than 1 page in size
   - Needs to be under the architectural offset limitations of loadstore
     scaled offsets. Which is 65KB for 128-bit vectors

Still keeps the pointer array around for cases when we would need
synthesize an address offset and it's just easier to load the
process-wide table.

The performance improvement here is removing the dependency in the
ldr+ldr chain. In microbenchmarks this has shown to have an improvement
of ~4% by removing this dependency chain on Cortex-X1C.
2023-10-04 20:56:29 -07:00
Ryan Houdek ee6debe8fd FEXCore: Adds DividePow2 helper 2023-10-04 20:56:29 -07:00
Mai 3ba1c7912c Merge pull request #3178 from Sonicadvance1/fix_avx_alias_precolour
Minor AVX optimizations
2023-10-04 21:31:20 -04:00
Ryan Houdek a408afaeb0 InstCountCI: Update for optimized AVX 2023-10-04 10:05:09 -07:00
Ryan Houdek fba7c4bedc IR/RA: Fixes register aliasing and pre-colouring for AVX
This is the cause of a bunch of redundant moves that shows up in
InstCountCI. Fixing this aliasing and pre-colouring issue causes a ton
of 256-bit operations to become optimal.
2023-10-04 10:04:06 -07:00
Ryan Houdek c52753e9c8 OpcodeDispatcher: Minor optimization in vzeroall
Using the cached zero value is less efficient than loading it in to the
register for all these cases.

Lets us use rename hardware more efficiently and removes a dependency
chain on a single register.

Original:
```
movi v2.2d, #0x0
mov z16.d, p7/m, z2.d
<... 16 more times>
mov z31.d, p7/m, z2.d
```

Result:
```
movi v16.2d, #0x0
<... 16 more times>
movi v31.2d, #0x0
```
2023-10-04 10:01:13 -07:00
Ryan Houdek e39634d314 Arm64: Fixes assert in VSQSHL/VSQSHR with SVE
When Dst != Vector then we need to pass Dst in to both Zd and Zdn.
Would have worked fine in a release build but assert build managed to
capture it.
2023-10-04 09:59:59 -07:00
Ryan Houdek 507cf82dad Merge pull request #3176 from neobrain/fix_thunks_unused_artifacts
Thunks: Only build guest target for libfex_thunk_test if FEXLinuxTests are enabled
2023-10-04 07:07:18 -07:00
Ryan Houdek 48fa4f1121 Merge pull request #3156 from neobrain/feature_thunk_data_layout_analysis
Thunks: Analyze data layout to detect platform differences
2023-10-04 07:06:49 -07:00
Tony Wasserka e06d609bf0 Thunks: Drop unused STRUCT_VERIFIER define from CMake 2023-10-03 11:43:29 +02:00
Tony Wasserka 0a09e04e33 Thunks: Only build guest target for libfex_thunk_test if FEXLinuxTests are enabled 2023-10-03 11:43:27 +02:00
Ryan Houdek a1a709f948 Merge pull request #3170 from Sonicadvance1/vixl_sim_instcountci
InstCountCI: Enable running on x86 hosts
2023-10-02 16:38:25 -07:00
Ryan Houdek 5925eef213 Github/InstCountCI: Enables x86 runner
To ensure we don't break this path for developers.
2023-10-02 16:26:14 -07:00
Ryan Houdek df369bd6a0 InstCountCI: Enable running on x86 hosts
This is a quality of life improvement for people that want to tinker
with the InstCountCI but they may not necessarily have an Arm64 device
available immediately for poking.

As long as the vixl disassembler is enabled then the InstCountCI tests
can run and get bit-accurate encodings just like on an Arm64 device.

This also ensures that behaviour is consistent with or without the vixl
simulator enabled which is very important when running on x86 hosts.
2023-10-02 16:26:14 -07:00
Ryan Houdek 978489fce1 InstCountCI: Explicitly disable SVE256 for one test group
These instructions are specifically testing the SVE128 implementations,
don't want SVE256 mucking up the instructions.
2023-10-02 16:26:14 -07:00
Ryan Houdek d5a4d9b17f InstCountCI: Adds option to disable cssc for tests
One x87 instruction was using CSSC abs
2023-10-02 16:26:14 -07:00
Ryan Houdek 9933ef07ea Tools: Enable indirect vixl runtime calls if simulator is used
So tests can still run.
2023-10-02 16:26:14 -07:00
Ryan Houdek 6964e65660 HostFeatures: Hardcode icache and dcache line size on x86
64-byte is effectively part of x86's ABI anyway. No need to query it for
our uses.
2023-10-02 16:26:14 -07:00
Ryan Houdek 11db8e7506 FEXCore: Wire up the new option to disable vixl indirect runtimes
Also so it compiles without the vixl simulator enabled.
2023-10-02 16:26:12 -07:00
Ryan Houdek b6b5e93dbb Config: Adds an option to disable vixl sim indirect runtime calls 2023-10-02 16:23:11 -07:00
Ryan Houdek 935b3a313a Merge pull request #3171 from Sonicadvance1/merge_dispatcher
FEXCore: Merge Arm64Dispatcher in to Dispatcher
2023-10-02 16:22:36 -07:00
Tony Wasserka fe681ab335 unittests/ThunkLibs: Specify clang resource directory when compiling test code 2023-10-02 22:18:23 +02:00
Tony Wasserka 2d9e816ff5 unittests/ThunkLibs: Add various tests for structs repacking and for void parameters 2023-10-02 22:18:23 +02:00
Tony Wasserka b04b0549a9 unittests/ThunkLibs: Add data layout tests 2023-10-02 22:18:22 +02:00
Tony Wasserka 2b472cb962 Thunks/gen: Enforce type compatibility for function parameters 2023-10-02 22:18:22 +02:00
Tony Wasserka 7f931b5623 Thunks/gen: Add detection logic for data layout differences
This runs the data layout analysis pass added in the previous change twice:
Once for the host architecture and once for the guest architecture. This
allows the new DataLayoutCompareAction to query architecture differences for
each type, which can then be used to instruct code generation accordingly.

Currently, type compatibility is classified into 3 categories:
* Fully compatible (same size/alignment for the type itself and any members)
* Repackable (incompatibility can be resolved with emission of automatable
  repacking code, e.g. when struct members are located at differing offsets
  due to padding bytes)
* Incompatible
2023-10-02 22:18:22 +02:00
Tony Wasserka 070fa9f924 Thunks/gen: Add data layout analysis
This adds a ComputeDataLayout function that maps a set of clang::Types
to an internal representation of their data layout (size, member list, ...).
2023-10-02 22:18:22 +02:00
Tony Wasserka 371bf50c76 Thunks/gen: Track data types passed across architecture boundaries
The set of these types is tracked in AnalysisAction, to which extensive
verification logic is added to detect potential incompatibilities and to
enforce use of annotatations where needed.
2023-10-02 22:18:22 +02:00
Tony Wasserka d65d29903b Thunks/gen: Rename EmitOutput to OnAnalysisComplete 2023-10-02 22:03:10 +02:00
Tony Wasserka 7791e0090d Thunks: Disable 32-bit host thunks
These are not supported yet.
2023-10-02 22:03:10 +02:00
Alyssa Rosenzweig 02da6d6ce7 Merge pull request #3174 from Sonicadvance1/remove_steam_appconfig
AppConfig: Removes Steam config
2023-10-01 18:48:30 -04:00
Ryan Houdek a478cbb694 AppConfig: Removes Steam config
This was only required on x86 devices trying to escape the emulation.
Since x86 is now remove, this is entirely unnecessary.

When Steam launches applications with `/bin/sh`, this will remain under
the emulation and not escape these days.
2023-10-01 08:46:53 -07:00
Ryan Houdek 3a25dd6d2b Merge pull request #3173 from CallumDev/x87f64-fabs
X87F64: Implement FABS with vector instruction
2023-10-01 01:54:11 -07:00
CallumDev 9c25db83d9 JIT: VectorOps remove extraneous element size logs 2023-10-01 15:03:21 +10:30
CallumDev 7346476546 Update InstCountCI 2023-10-01 14:41:13 +10:30
CallumDev c42b581378 X87F64: Implement FABS with vector instruction 2023-10-01 14:39:55 +10:30
Ryan Houdek ccfd770d9d Merge pull request #3172 from CallumDev/x87f64-opts
X87F64: Use Bfe for rounding mode, FCHS use float instruction
2023-09-30 18:41:29 -07:00
CallumDev d4a623a3fb InstCountCI Update 2023-10-01 11:22:18 +10:30
CallumDev c09c25005e X87F64: Use Bfe for rounding mode, FCHS use float instruction 2023-10-01 11:11:33 +10:30
Ryan Houdek 90570fd5f4 FEXCore: Merge Arm64Dispatcher in to Dispatcher
With the removal of the x86 JIT, there is no need to have these be
independent classes.

Merges the Arm64Dispatcher in to the base Dispatcher class.
No functional change, just moving code.
2023-09-30 09:31:55 -07:00
Mai ab4642af38 Merge pull request #3167 from Sonicadvance1/gatherqdps
unittests/ASM: Implements tests for vpgatherqd/vgatherqps
2023-09-29 12:16:43 -04:00
Mai d94e5ce7f4 Merge pull request #3168 from Sonicadvance1/gatherqqpd
unittests/ASM: Implements tests for vpgatherqq/vgatherqpd
2023-09-29 12:16:12 -04:00
Mai dad7086fd0 Merge pull request #3166 from Sonicadvance1/gatherdqpd
unittests/ASM: Implements tests for vpgatherdq/vgatherpq
2023-09-29 12:15:39 -04:00
Ryan Houdek a21def7d74 unittests/ASM: Implements tests for vpgatherqq/vgatherqpd
Similar to previous tests, vpgatherqq and vgatherqpd are equivalent
instructions. So the tests are the same with the mnemonic changed.

This adds tests for an additional two sets of instructions. Getting us
full coverage of all eight instructions if we include the tests from
PR #3167 and #3166

Tests the same things as described in #3165

In addition, since these tests use 64-bit indices for address
calculation, we can easily generate and indice vector that tests
overflow. So every test at every displacement ALSO gains an additional
overflow test to ensure correct behaviour around pointer overflow
calculation.
2023-09-29 08:04:47 -07:00
Ryan Houdek 0d8d5444a4 unittests/ASM: Implements tests for vpgatherqd/vgatherqps
Similar to previous tests, vgatherqd and vgatherqps are equivalent
instructions. So the tests are the same with the mnemonic changed.

This adds tests for an additional two sets of instructions, Getting us
up to six total over the eight if we include the tests from #3166.

Tests the same things as described in #3165

In addition, since these tests use 64-bit indices for address
calculation, we can easily generate and indice vector that tests
overflow. So every test at every displacement ALSO gains and additional
overflow test to ensure correct behaviour around pointer overflow
calculation.
2023-09-29 07:20:07 -07:00
Ryan Houdek eedfad5036 unittests/ASM: Implements tests for vpgatherdq/vgatherpq
Just like the previous tests, vpgatherdq and vgatherpq are equivalent
instructions. So the tests are the same except for the instruction
mnemonic again.

This adds unittests for two more of the eight gather instructions.
Getting us up to testing four in total.
Specifically this adds tests for 32-bit indices while loading 64-bit
element instructions.

Same thing as PR #3165 for what it tests versus doesn't.
2023-09-28 22:49:03 -07:00
Ryan Houdek 85da0f0640 Merge pull request #3165 from Sonicadvance1/gatherddps
unittests/ASM: Implements tests for vpgatherdd/vgatherps
2023-09-28 22:44:38 -07:00
Ryan Houdek 9a01b440e3 unittests/ASM: Implements tests for vpgatherdd/vgatherps
vpgatherdd and vgatherps are effectively the same instructions, so the
tests are the same except for the instruction mnemonic.

This adds unit tests for two of the eight gather instructions.
Specifically this adds tests for the 32-bit indices loading 32-bit
elements instructions.

What it tests:
- Tests all displacement scales
- Tests multiple mask arrangements
- Ensures the mask register is zero'd after the instruction

What it doesn't test:
- Doesn't test address size calculation overflow
   - Only would happen on 32-bit with 32-bit indices, or /really/ high
     base addresses
   - The instruction should behave as a mask to the address size
   - Effectively behaves like `(uint64_t)(base + index << ilog2(scale))`
   - Better idea is to just not expose AVX to 32-bit applications
- Doesn't test VSIB immediate displacement
   - This just ends up being base_addr + imm so it isn't too interesting
   - We can add more tests in the future if we think we messed that up
- Doesn't test partial fault behaviour
   - Because that's a nightmare.

Specifically keeps each instruction test small and isolated so if a
single register fails it is very easily to nail down which operation did
it.
I know some of our ASM tests do a chunk of work and spit out a result at
the end which can be difficult to debug in some cases. Didn't want to do
that which is why the tests are spread out across 16 files for these
single class of instructions.
2023-09-28 19:58:34 -07:00
Ryan Houdek 228ee7fa47 TestHarnessRunner: Support AVX2 flag detection 2023-09-28 19:58:34 -07:00
Ryan Houdek 98789a8039 FEXCore: Implement support for AVX2 feature detection 2023-09-28 19:57:08 -07:00
Ryan Houdek 14398742c3 Merge pull request #3164 from neobrain/fix_thunks_asan
Thunks: Fix AddressSanitizer build
2023-09-28 12:05:55 -07:00
Tony Wasserka 5a7e3192da Thunks: Fix AddressSanitizer build 2023-09-28 15:13:03 +02:00
Ryan Houdek 6b4ff4ae81 Merge pull request #3163 from alyssarosenzweig/opt/ascii-flags
Optimize ASCII flags
2023-09-27 10:42:47 -07:00
Ryan Houdek d1d3de80d1 Merge pull request #3157 from alyssarosenzweig/opt/unmask-in
OpcodeDispatcher: Don't mask logic op inputs
2023-09-27 10:38:12 -07:00
Alyssa Rosenzweig 2e32e1367d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:55:57 -04:00
Alyssa Rosenzweig 711583aa76 OpcodeDispatcher: Optimize PTEST flags
Zero NZCV first to avoid RMW.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:55:57 -04:00
Alyssa Rosenzweig 3efac9646c OpcodeDispatcher: Optimize ASCII flags
Make the zeroing of undefined NZCV more obvious. Mitigates regressions from
future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:31:31 -04:00
Alyssa Rosenzweig 095a362046 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:30:09 -04:00
Alyssa Rosenzweig 3bb64c64e3 OpcodeDispatcher: Don't mask for TEST
Like AND.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:30:02 -04:00
Alyssa Rosenzweig a4de164944 OpcodeDispatcher: Use lshr for ah/bh with AllowUpperGarbage
If we ever get around to fusing ops with shifts in the ConstProp optimizer (may
or may not be worthwhile), this will delete an instruction from things like "or
al, bh".

Even though lsr is the same speed as bfe on Firestorm, I feel if you ask for
garbage you should get garbage C:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:28:01 -04:00
Alyssa Rosenzweig 45a645fbbc OpcodeDispatcher: Don't mask logic op inputs
Pointless, upper bits ignored anyway. Deletes piles of uxt and even some 32-bit
instruction moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:12:22 -04:00
Alyssa Rosenzweig 92211bf8c6 OpcodeDispatcher: Add AllowUpperGarbage option
To load 8-bit sources without bfe'ing for al/bl/cl if the caller knows it
doesn't need masking behaviour, but without lying about the size so the extract
for ah/bh/ch will still work properly.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:08:20 -04:00
Alyssa Rosenzweig 728d3f8ac7 InstCountCI: Add a case with a hi 8-bit reg
Noticeably different code pattern.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 18:33:55 -04:00
Ryan Houdek ca87d8688d Merge pull request #3153 from alyssarosenzweig/opt/adcs
Use adcs
2023-09-26 09:57:01 -07:00
Ryan Houdek e32601f49d Merge pull request #3161 from neobrain/fix_ctest_silent_failures
unittests: Instruct CTest to print output from tests on failure
2023-09-26 08:26:15 -07:00
Tony Wasserka f4dd456c80 unittests: Instruct CTest to print output from tests on failure 2023-09-26 17:16:28 +02:00
Alyssa Rosenzweig 7b22dbfe24 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 10:05:59 -04:00
Alyssa Rosenzweig 7a06cc9727 IR: Use adcs/sbcs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 09:06:46 -04:00
Ryan Houdek 8b3881b5db Merge pull request #3154 from alyssarosenzweig/opt/smol-carry
Optimize 8/16-bit CF calculation
2023-09-26 05:49:07 -07:00
Ryan Houdek 76d4637d9c Merge pull request #3159 from neobrain/feature_update_vulkan
Thunks: Update Vulkan thunk to v1.3.261.1
2023-09-26 05:20:18 -07:00
Alyssa Rosenzweig 0d12cce74f Merge pull request #3158 from Sonicadvance1/unittest_for_3153
unittests/ASM: Adds unit test caught by #3153
2023-09-26 08:15:40 -04:00
Tony Wasserka 04592af609 Thunks: Update Vulkan thunk to v1.3.261.1 2023-09-26 12:14:58 +02:00
Ryan Houdek d8366c04dc unittests/ASM: Adds unit test caught by #3153 2023-09-26 00:28:45 -07:00
Ryan Houdek 533f35934c Merge pull request #3155 from neobrain/opt_thunks_rebuilds
Thunks: Avoid recompiling thunk interfaces on FEXLoader changes
2023-09-25 19:21:09 -07:00
Alyssa Rosenzweig 35bb7cc801 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 19:41:31 -04:00
Alyssa Rosenzweig 5facb21d30 OpcodeDispatcher: Don't mask small add/sub carries
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.

Deletes a pile of code both in FEX and the output :-)

ADC/SBC could probably get similar treatment later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 18:25:30 -04:00
Tony Wasserka adead832a5 Thunks: Avoid recompiling thunk interfaces on FEXLoader changes
The interface files themselves don't use FEXLoader. Only the final library
does.
2023-09-25 23:04:09 +02:00
Ryan Houdek 5eed24a242 Merge pull request #3152 from Sonicadvance1/instcountci_x87_f64
InstCountCI: Support f64 reduced precision mode tests
2023-09-24 19:29:37 -07:00
Ryan Houdek 7907f70ed2 InstCountCI: Adds new x87 reduced precision mode tests 2023-09-24 18:50:05 -07:00
Ryan Houdek 7141332f6f InstCountCI: Support setting environment variables in tests
This will allow us to enable FEX options through environment variables
just like the ASM tests.
2023-09-24 18:50:01 -07:00
Ryan Houdek 234e029391 Merge pull request #3145 from Sonicadvance1/optimize_inline_calls
PassManager: Optimize out CPUID and XGetBV calls
2023-09-24 18:09:18 -07:00
Ryan Houdek 19a7b514e6 Merge pull request #3150 from alyssarosenzweig/opt/ornror
Optimize PF calculation in lahf
2023-09-24 18:05:57 -07:00
Ryan Houdek 220761a0e8 Merge pull request #3151 from Sonicadvance1/unique_name_workflow_jobs
Github: Changes jobs to have unique names
2023-09-24 18:03:57 -07:00
Alyssa Rosenzweig cbd4daddff InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig c8519b0b87 OpcodeDispatcher: Remove LoadPF
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig 68d32ad70d OpcodeDispatcher: Optimize PF in lahf
Use the raw popcount rather than the final PF and use some sneaky bit math to
come out 1 instruction ahead.

Closes #3117

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Ryan Houdek 62890f148f Github: Changes jobs to have unique names
These overlapping names make it impossible to ensure all checks are
required to pass before merge.

Unique names will fix this.
2023-09-24 17:52:47 -07:00
Alyssa Rosenzweig 1f02a6da34 IR: Add Ornror op
Mostly copypaste of Orlshl... we really should deduplicate this mess somehow.
Maybe a shift enum on the core Or op?

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:47:50 -04:00
Alyssa Rosenzweig 86063411dc Revert "OpcodeDispatcher: Use plain Lshl for flags"
This logic is unused since 8adfaa9aa ("OpcodeDispatcher: Use SelectCC for x87"),
which addressed the underlying issue.

This reverts commit df3833edbe.
2023-09-24 20:47:50 -04:00
Ryan Houdek 9968e6431f Passes: Rename SyscallOptimization
This is now inlining multiple external calls out of the JIT. Rename it
to InlineCallOptimization.
2023-09-24 17:25:38 -07:00
Ryan Houdek ff24f64b2a PassManager: Optimize out CPUID and XGetBV calls
If we const-prop the required functions and leafs then we can directly
encode the CPUID information rather than jumping out of the JIT.
In testing almost all CPUID executions const-prop which function is
getting called. Worst case that I found was only 85% const-prop rate.

This isn't quite 100% optimal since we need to call the RCLSE and
Constprop passes after we optimize these, which would remove some
redundant moves.

Sadly there seems to be a bug in the constprop pass that starts crashing
applications if that is done.
Easily enough tested by running Half-Life 2 and it immediately hitting
SIGILL.

Even without this optimization, this is stil a significant savings since
we aren't jumping out of the JIT anymore for these optimized CPUIDs.
2023-09-24 17:25:38 -07:00
Ryan Houdek e9a7ef2534 CPUID: Describe CPUID functions if they return constant state or not
Most CPUID routines return constant data, there are four that don't.
Some CPUID functions also need the leaf descriptor, so we need to
describe that as well.

Functions that don't return constant data:
- function 1Ah - Returns different data depending on current CPU core
- function 8000_000{2,3,4} - Different data based on CPU core

Functions that need leaf constprop:
- 4h, 7h, Dh, 4000_0001h, 8000_001Dh
2023-09-24 17:25:38 -07:00
Ryan Houdek 842c57e221 CPUID: Constify some functions
These don't modify CPUIDEmu state.
2023-09-24 17:25:38 -07:00
Ryan Houdek 93aeb157b4 Merge pull request #3149 from Sonicadvance1/fail_on_change
InstCountCI: Fail CI if there was any difference.
2023-09-24 17:23:52 -07:00
Ryan Houdek 02ff9f200c InstCountCI: Upload diff and check for failure 2023-09-24 17:14:08 -07:00
Ryan Houdek f65b40f298 InstCountCI: Fail if inst count has changed 2023-09-24 17:14:08 -07:00
Ryan Houdek c38beff826 Merge pull request #3148 from Sonicadvance1/add_negative_primaries
InstCountCI: Adds negative immediate primary tests
2023-09-24 17:13:34 -07:00
Ryan Houdek 94c22b2269 InstCountCI: Adds negative immediate primary tests
Noticed these were missing
2023-09-24 17:02:58 -07:00
Ryan Houdek bee97309f6 Merge pull request #3147 from alyssarosenzweig/opt/0924
More opts to the dispatcher + 1 to the JIT
2023-09-24 17:01:37 -07:00
Alyssa Rosenzweig 331941dec6 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig 8798e0cba0 Arm64: Rewrite Set/GetRoundingMode
I went auditing for places to use cset and what I found was hot garbage.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig c5fc03dac4 OpcodeDispatcher: Use cset for blsr/etc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig e63871ed2e OpcodeDispatcher: Handle sub in CalculateOF
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig ea8b7633eb OpcodeDispatcher: Optimize OF calc of immediates
If we know the sign of one of the sources, we can do better when calculating OF.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 18:16:09 -04:00
Ryan Houdek e795ec683d Merge pull request #3139 from Sonicadvance1/workaround
FEXServerClient: Adds back ServerSocketPath config option
2023-09-23 17:09:04 -07:00
Ryan Houdek 6dc5c0d3be Merge pull request #3144 from Sonicadvance1/optimize_redundant_store_load
RCLSE: Optimize redundant store->load operations
2023-09-23 17:06:10 -07:00
Ryan Houdek eb5e0be569 FEXServerClient: Adds back ServerSocketPath config option
This option was disabled a few months ago when we switched the server
socket from a filesystem unix socket to an abstract socket.
This partially broke our chroot scripts which relied on this option
existing.

Readds support for an explicitly named abstract socket named from
config.

This is a workaround for dealing with chroots that change users.
They end up changing a user while doing operations and then can't
connect to the FEXServer anymore because environment variables have been
wiped away.
2023-09-23 16:59:58 -07:00
Ryan Houdek be3ff804a6 InstCountCI: Update for optimization 2023-09-23 06:11:35 -07:00
Ryan Houdek 9ab2967d71 Arm64: Fixes wide shifts
movprfx is invalid to use when the source register matches the movprfx
destination.

This was getting picked up on by `TwoByte/0F_D1.asm` now that RCLSE is
working better now.
2023-09-23 06:06:18 -07:00
Ryan Houdek d01b457727 RCLSE: Optimize redundant store->load operations
The bug that was causing crashes with this was due to inline syscalls.
Now that this is fixed we can re-enable store->load operations.

This allows constant propagation to work significantly better, which
means inline syscalls start working again. This can significantly
improve syscall performance in some cases.

This is most likely to improve performance in dxsetup and vc_redist but
hard to get a real profile.

Additionally this will let us inline cpuid results in the future which
is pretty nice.
2023-09-23 06:06:18 -07:00
Mai 4e9a114858 Merge pull request #3142 from Sonicadvance1/inline_syscall_fix
Arm64: Fixes inline syscalls
2023-09-23 09:03:49 -04:00
Mai 72d092e951 Merge pull request #3141 from Sonicadvance1/fix_simm9_range
ConstProp: Fixes unscaled signed 9-bit range
2023-09-23 09:03:01 -04:00
Mai da3e172857 Merge pull request #3140 from Sonicadvance1/fix_core_sanitization
Config: Fixes core sanitization
2023-09-23 09:01:42 -04:00
Ryan Houdek 28fa0bda31 Arm64: Fixes inline syscalls
Ever since we reordered registers in `X86Enums.h` this has silently been
broken. This wasn't hit because RCLSE has been broken ever since SRA was
added, so inlinesyscalls just weren't ever happening.

Quick fix while I think of a way to more strictly correlate these
registers so it doesn't happen again.
2023-09-23 02:56:32 -07:00
Ryan Houdek 1f2a3cfa8b ConstProp: Fixes unscaled signed 9-bit range
The range was slightly incorrect which mostly wouldn't have caused
issues.

The lowest byte would have just generated slightly less optimal code.
The upper byte could have generated broken code, which our CI couldn't
catch since TSO instructions only get enabled when multiple threads are
in-flight.

Easy enough to fix.
2023-09-23 01:13:54 -07:00
Ryan Houdek 571b0fe47e Config: Fixes core sanitization
This would have caused core to try and initialize a custom core on
Arm64, which causes a std::function assert because it doesn't support
that.

Users would likely get hit by this immediately since we deleted the
interpreter and shifted all the core numbers.
2023-09-23 00:52:23 -07:00
Ryan Houdek 86ad35c418 Merge pull request #3138 from alyssarosenzweig/opt/train
Requiem for the x86 jit
2023-09-22 16:33:15 -07:00
Alyssa Rosenzweig 0b27029c3f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:10:41 -04:00
Alyssa Rosenzweig 223a6562ff IR: Support <32-bit TestNZ
Originally this was going to use setf8/setf16, but it looks like the approach of
shift-and-test turns out to be faster. As a bonus this is a nice delete-the-code
win :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:08:26 -04:00
Alyssa Rosenzweig b1231c24ef OpcodeDispatcher: Omit AF xor for common constants
The only reason we need to XOR arguments for AF is to get bit 4 correct. But if
the operand in question is known to have bit 4 clear, the XOR will be an
effective no-op and can be skipped. This saves an instruction in a bunch of
common cases, like inc/dec. If we dedicated a register to AF to eliminate the
store, we would not save an instruction from this but would still come out ahead
due to an eor turning into a (zero cycle?) mov that can be handled by the
renamer.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:08:26 -04:00
Alyssa Rosenzweig 699aa85c4b OpcodeDispatcher: Opt PF selection
Fold the and in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Alyssa Rosenzweig 2d65a3677b OpcodeDispatcher: Optimize NZCV selects
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Alyssa Rosenzweig 2a2619c0f5 IR: Add bit masking selects
Add new synthetic condition codes that do an AND as their relational operator,
testing the result. This is 1 IR op for things like

  (A & B) == 0 ? C : D

This can translate to

  tst A, B
  csel A, B, eq

In the future, if A is the NZCV register and B is a supported immediate, eg

  (NZCV & 0x80000000) == 0 ? C : D

this will be able to translate to a single instruction with the appropriate
condition

  csel A, B, pl

but that needs RA support.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Ryan Houdek 797c890ff6 Merge pull request #2874 from bylaws/wowfex
Add WOW64 JIT frontend
2023-09-22 15:47:59 -07:00
Ryan Houdek 879b41c184 Merge pull request #3134 from Sonicadvance1/remove_x86_jit
FEXCore: Removes x86 JIT.
2023-09-22 15:36:47 -07:00
Ryan Houdek 0fbf403787 Adds back in host testharnessrunner CI
Necessary for asm tests to still run in the host "core".
Useful for ensuring correct behaviour of our assembly tests.
2023-09-22 14:46:03 -07:00
Billy Laws 04cf418452 Windows: Add SPDX license identifiers 2023-09-22 10:12:40 -07:00
Billy Laws 057a7c6ee8 WOW64: Implement thread suspension handling
This provides more robust handling than a signal based approach, as the
suspender is able to wait for the suspendee to reach a suitable position and
flush its context to memory before returning.
2023-09-22 10:12:40 -07:00
Billy Laws 3d6955592b WOW64: Implement partial self-modifying code handling
This should support most simple cases of SMC, however programs which make use
of separate shared memory mappings for writing and execution are not handled.
The overall approach is the same as is done for linux, where RWX mappings are
protected to RX and then when a write occurs the signal handler invalidates the
faulting page and reprotects it to RWX until code in that page is jitted again.
2023-09-22 10:12:40 -07:00
Billy Laws f57aee0a62 WOW64: Add a templated interval list implementation
Stores binary intervals in a sorted vector container, to be used for SMC
handling.
2023-09-22 10:12:40 -07:00
Billy Laws c978fdd12f WOW64: Implement basic code invalidation handling 2023-09-22 10:12:40 -07:00
Billy Laws 19713bd20a WOW64: Implement exception handling with context restoration
When an exception occurs, pretend that we were just at the point of JIT entry
so the stack can be unwound to the wow64 SEH handler, which then handles
dispatching the exception to the x86 guest with the restored context.
2023-09-22 10:12:40 -07:00
Billy Laws 22b1fea96d WOW64: Handle unaligned atomic accesses
This is done in EnsureConsistentState rather than as a VEH to avoid needing to
go through all of wine's exception handling logic for such a hot path.
2023-09-22 10:12:40 -07:00
Billy Laws be4fcaf65c WOW64: Report CPU features based off of the emulated cpuid 2023-09-22 10:12:40 -07:00
Billy Laws 2add8a7751 Windows: Introduce a barebones FEXCore-based WOW64 BT module
This allows for running x86 applications under wine without having to run all
of wine under FEX. The JIT is invoked when running application code and then
left when handling NT syscalls or unix calls to e.g. the Vulkan driver.
2023-09-22 10:12:40 -07:00
Billy Laws 9612133088 Windows: Generate import libraries for private ntdll and wow64 APIs
The MinGW supplied import libraries are incomplete and miss a lot of
functions necessary to implement lower level windows code. To avoid
needing to many resolve every function, pull in .def files from wine
that detail the entire ntdll and wow64 APIs.
2023-09-22 10:12:40 -07:00
Billy Laws f46fd42977 Windows: Add a minimal set of wine-derived headers
These are cut down versions of wine headers containing only what is necessary
for WOW. This shouldn't carry any license implications for FEX, as per the
LGPLv3 license:

```
The object code form of an Application may incorporate material from a header
file that is part of the Library. You may convey such object code under terms
of your choice, provided that, if the incorporated material is not limited to
numerical parameters, data structure layouts and accessors, or small macros,
inline functions and templates (ten or fewer lines in length), you do both of
the following:

a) Give prominent notice with each copy of the object code that the Library is
used in it and that the Library and its use are covered by this License.
b) Accompany the object code with a copy of the GNU GPL and this license
document.
```
2023-09-22 10:12:40 -07:00
Billy Laws 51f8c83c76 Context: Add an alternative thread-oriented execute function 2023-09-22 10:12:40 -07:00
Billy Laws d641d3f61e OpcodeDispatcher: Avoid redundantly passing args to WIN32 ABI syscalls 2023-09-22 10:12:39 -07:00
Ryan Houdek 02ae59a348 github: Disables default build test on x64 2023-09-21 18:30:03 -07:00
Ryan Houdek 64df9e31c6 github: Remove mingw tests from x86 CI 2023-09-21 18:30:03 -07:00
Ryan Houdek d32bb993a8 github: Remove glibc fault tests from x86 CI 2023-09-21 18:30:03 -07:00
Ryan Houdek b5cc9a12f2 FEXCore: Removes x86 JIT.
This is blocking performance improvements. This backend is almost
unilaterally unused except for when I'm testing if games run on Radeon
video drivers.

Hopefully AmpereOne and Orin/Grace can fulfill this role when they
launch next year.
2023-09-21 18:30:02 -07:00
Ryan Houdek 65b6df9dbb Merge pull request #3133 from Sonicadvance1/remove_vestigial_interpreter
FEXCore: Removes vestigial Interpreter code
2023-09-21 18:15:32 -07:00
Ryan Houdek 31564354b1 FEXCore: Removes vestigial Interpreter code 2023-09-21 15:49:49 -07:00
Ryan Houdek fea72ce19c Merge pull request #3120 from Sonicadvance1/more_optimal_x87
FEXCore: Support preserve_all ABI for interpreter fallbacks
2023-09-21 15:35:37 -07:00
Ryan Houdek 2b7e1d10ec Merge pull request #3131 from Sonicadvance1/optimize_btr
OpcodeDispatcher: Optimize lock btr
2023-09-21 15:06:55 -07:00
Ryan Houdek 5444810d64 Merge pull request #3132 from alyssarosenzweig/opt/orlshl
Optimize reconstructing x87, harder
2023-09-21 15:02:37 -07:00
Ryan Houdek 4a2ceabfdd InstCountCI: Add atomic bit test instructions
These all can likely be more optimal.
2023-09-21 14:54:51 -07:00
Ryan Houdek 1a4d1d820b OpcodeDispatcher: Optimize lock btr
This is an atomicFetchCLR, removes two mvn instructions that are back to
back negating the source.

We didn't have this instruction combination in InstCountCI so will be a
bit hard to see.
2023-09-21 14:54:51 -07:00
Ryan Houdek 0ae4bbb9c5 IR: Implements support for AtomicFetchCLR
This is the native ARM operation rather than fetchAnd. Will make an
instruction an instruction slightly more optimal.
2023-09-21 14:54:51 -07:00
Ryan Houdek 7d99eb05c6 Merge pull request #3128 from alyssarosenzweig/rm/interp
FEXCore: Gut interpreter
2023-09-21 14:51:44 -07:00
Alyssa Rosenzweig 8247ded2cf unittests: Remove stale comments
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 12:48:12 -04:00
Alyssa Rosenzweig c52741c813 FEXCore: Gut interpreter
It is scarcely used today, and like the x86 jit, it is a significant
maintainence burden complicating work on FEXCore and arm64 optimization. Remove
it, bringing us down to 2 backends.

1 down, 1 to go.

Some interpreter scaffolding remains for x87 fallbacks. That is not a problem
here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 12:48:12 -04:00
Alyssa Rosenzweig 75ffbc16f2 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:14:08 -04:00
Alyssa Rosenzweig 1596e33f58 OpcodeDispatcher: Remove pointless or
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig 07d03f1610 OpcodeDispatcher: Don't opencode bfe, badly
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig a8b48dcacd OpcodeDispatcher: Swap some selects
...if it lets us use cset.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig bb87b2a19d OpcodeDispatcher: Use more Orlshl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig 19eff62c77 OpcodeDispatcher: Use orlshl for FCW
Potentially easier on the RA (bfi has a tied operand), mostly whatever here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 08:55:25 -04:00
Mai 5fc8699db9 Merge pull request #3130 from Sonicadvance1/optimize_fsw
OpcodeDispatcher: Optimize reconstructing FSW
2023-09-21 08:35:16 -04:00
Mai 43fd159689 Merge pull request #3129 from Sonicadvance1/remove_non_explicit_selectcc
OpcodeDispatcher: Removes non-explicit SelectCC function
2023-09-21 08:33:30 -04:00
Ryan Houdek 758820ca86 InstCountCI: Update for optimized FSW reconstruction 2023-09-21 02:27:04 -07:00
Ryan Houdek 5664195e49 OpcodeDispatcher: Optimize reconstructing FSW
Minor optimization using Bfi to insert C0, C1, C2, & C3
2023-09-21 02:07:27 -07:00
Ryan Houdek 683daefc15 InstCountCI: Minor changes 2023-09-21 01:57:08 -07:00
Ryan Houdek 8e9e87f631 OpcodeDispatcher: Removes non-explicit SelectCC function
Renames the explicit sized one to `SelectCC`
Cleans up a bit of duplicated code.
2023-09-21 01:56:38 -07:00
Ryan Houdek 0a0865eb1c InstCountCI: Update for minor change 2023-09-20 18:51:18 -07:00
Ryan Houdek d588d41ab9 InterpreterFallbacks: Converts X87 and String ops to preserve_all
This improves performance!
2023-09-20 18:51:18 -07:00
Ryan Houdek 8aa8d597f6 Arm64: Supports jumping out of the JIT with preserve_all ABI
This improves perferformance when jumping out of the Arm64 JIT by
reducing the number of registers we need to save.
2023-09-20 18:51:18 -07:00
Ryan Houdek 67680d71a4 Merge pull request #3125 from Sonicadvance1/spdx_fexcore
FEXCore: Adds SPDX identifier
2023-09-19 17:42:07 -07:00
Ryan Houdek d86f41e29a Merge pull request #3124 from Sonicadvance1/spdx_fexcore_include
FEXCore/Include: Adds SPDX identifier
2023-09-19 17:41:59 -07:00
Ryan Houdek ba56e514bd Merge pull request #3123 from Sonicadvance1/spdx_fex_linux
FEX: Moves Linux utils and adds spdx
2023-09-19 17:41:51 -07:00
Ryan Houdek 9f5f09b772 Merge pull request #3122 from Sonicadvance1/spdx_fex_common
FEX/Common: Adds SPDX identifier
2023-09-19 17:41:44 -07:00
Ryan Houdek ddf4b5cbd4 Merge pull request #3121 from Sonicadvance1/spdx_tools
FEX/Tools: Adds SPDX identifier
2023-09-19 17:41:34 -07:00
Ryan Houdek e4613477b1 FEXCore/Interface/Core: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek d18ce59187 FEXCore/Interface/Core/JIT: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 1032224d62 FEXCore/Interface/Core/Interpreter: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 44767901fe FEXCore/Interface/Core/Dispatcher: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 1220c86573 FEXCore/Interface/Core/ArchHelpers: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 6ace406a2f FEXCore/Interface/Core/ObjectCache: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 38f1536255 FEXCore/Interface/Core/OpcodeDispatcher: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 324473651e FEXCore/Interface/Core/X86Tables: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 573148b27a FEXCore/Interface/Core/VSyscall: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek 12e1c2eaa0 FEXCore/Interface/Context: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek c678ea3060 FEXCore/Interface/GDBJIT: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Ryan Houdek e570b07ba0 FEXCore/Interface/Thunks: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek afda4d6b7a FEXCore/Interface/IR: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 22daa506f6 FEXCore/Interface/Config: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek e85b90c614 FEXCore/Common: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 9d3d33fa27 FEXCore/Utils: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 0d9dce987d Merge pull request #3126 from neobrain/feature_better_wayland_thunks64
Thunks/wayland: Add support for APIs required by zink and Super Meat Boy
2023-09-19 10:44:53 -07:00
Ryan Houdek 65d558b2c4 Merge pull request #3119 from alyssarosenzweig/opt/x87-sel
Make x87 FCMOV slightly less terrible
2023-09-19 10:34:29 -07:00
Tony Wasserka b00d413961 Thunks/wayland: Add more message signatures required by Super Meat Boy with zink 2023-09-19 17:33:24 +02:00
Tony Wasserka 6b54540756 Thunks/wayland: Add support for message signatures with nullable arguments 2023-09-19 17:33:24 +02:00
Tony Wasserka 356a42d330 Thunks/wayland: Reorder listener signatures alphabetically 2023-09-19 17:33:24 +02:00
Tony Wasserka 8fcf419183 Thunks/wayland: Add more functions required by Super Meat Boy via libdecor and SDL 2023-09-19 17:33:23 +02:00
Alyssa Rosenzweig 83c8b64c50 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig 25943d1d17 OpcodeDispatcher: Sigh.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig bf03dab295 Arm64: Use csetm
Saves some moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig 8adfaa9aa6 OpcodeDispatcher: Use SelectCC for x87
Better code gen and will benefit from future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:37:54 -04:00
Ryan Houdek ca6570d5de FEXCore/Include: Adds SPDX identifier 2023-09-18 22:13:10 -07:00
Ryan Houdek 3026f7249c FEX/Tools/CommonTools/Linux: Adds SPDX identifier 2023-09-18 22:03:29 -07:00
Ryan Houdek bea29fd2ba FEX: Moves some Linux utils to CommonTools
Was kind of in a weird place before.
2023-09-18 22:01:56 -07:00
Ryan Houdek fc55091fc5 FEX/Common: Adds SPDX identifier 2023-09-18 21:52:20 -07:00
Ryan Houdek 782cf3f7c7 Tools/Opt: Remove. Unused. 2023-09-18 21:45:25 -07:00
Ryan Houdek 8f25e9d3e6 FEXLoader: Adds SPDX identifier 2023-09-18 21:44:23 -07:00
Ryan Houdek 01175e2e7c FEXLoader/LinuxSyscalls: Adds SPDX identifier 2023-09-18 21:43:17 -07:00
Ryan Houdek 5d9d539495 FEXLoader/LinuxSyscalls/Utils: Adds SPDX identifier 2023-09-18 21:38:13 -07:00
Ryan Houdek efb5624db6 FEXLoader/LinuxSyscalls/EmulatedFiles: Adds SPDX identifier 2023-09-18 21:37:51 -07:00
Ryan Houdek fe0a16f478 FEXLoader/HostRunner: Adds SPDX identifier 2023-09-18 21:37:02 -07:00
Ryan Houdek d9d376d40d FEXLoader/ArchHelpers: Adds SPDX identifier 2023-09-18 21:36:23 -07:00
Ryan Houdek 74e7f88449 FEXLoader/AOT: Adds SPDX identifier 2023-09-18 21:35:59 -07:00
Ryan Houdek b2811ffc59 FEX/CommonGUI: Adds SPDX identifier 2023-09-18 21:35:25 -07:00
Ryan Houdek f08e1da577 FEX/CommonTools: Adds SPDX identifier 2023-09-18 21:35:07 -07:00
Ryan Houdek e863eba364 FEXGetConfig: Adds SPDX identifier 2023-09-18 21:34:42 -07:00
Ryan Houdek 10081595af FEXGDBReader: Adds SPDX identifier 2023-09-18 21:34:22 -07:00
Ryan Houdek 75d53725e5 FEXConfig: Adds SPDX identifier 2023-09-18 21:33:02 -07:00
Ryan Houdek d21335be85 FEXBash: Adds SPDX identifier 2023-09-18 21:32:25 -07:00
Ryan Houdek e0385cd807 FEXRootFSFetcher: Adds SPDX identifier 2023-09-18 21:31:51 -07:00
Ryan Houdek e962462e79 FEXServer: Adds SPDX identifier 2023-09-18 21:31:18 -07:00
Ryan Houdek 5896c30954 CodeSizeValidation: Adds SPDX identifier 2023-09-18 21:30:30 -07:00
Ryan Houdek 745729cdc2 SoftFloat-3e: Adds preserve_all attribute to all functions used
This will let FEX's JIT be more optimal
2023-09-18 17:42:48 -07:00
Ryan Houdek 95e5d37e4c FEXCore: Adds compile time check support for preserve_all 2023-09-18 17:09:54 -07:00
Ryan Houdek 838293c2f0 FEXCore: Remove unused FallbackhandlerIndex LoadFCW
We removed this once passing in FCW explicitly.
2023-09-18 17:06:46 -07:00
Ryan Houdek da21fc937b FHU: Fixes syscall helper caching
check_cxx_source_compiles caches by variable name, so `compiles` was
getting cached and breaking future checks.
2023-09-18 17:05:40 -07:00
Alyssa Rosenzweig 3b188b7f49 Merge pull request #3118 from Sonicadvance1/spdx_fhu
FHU: Prepend SPDX identifier
2023-09-18 18:48:30 -04:00
Ryan Houdek 94bbd415a2 FHU: Prepend SPDX identifier
Added with `sed -i '1 i\\/\/ SPDX-License-Identifier: MIT' *.h`
2023-09-18 11:45:18 -07:00
Ryan Houdek 2ea2300408 Merge pull request #3110 from Sonicadvance1/buffered_jit_symbols
FEXCore/JitSymbols: Buffer writes to reduce overhead
2023-09-18 11:38:06 -07:00
Ryan Houdek 000fb2efae Merge pull request #3068 from neobrain/feature_thunk_testlib
unittests: Add test thunk library
2023-09-18 10:31:42 -07:00
Ryan Houdek 8b523082af Merge pull request #3116 from alyssarosenzweig/minor/flag-opts
Minor/flag opts
2023-09-18 10:28:51 -07:00
Alyssa Rosenzweig 5d2a3cd322 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 11:01:46 -04:00
Alyssa Rosenzweig df3833edbe OpcodeDispatcher: Use plain Lshl for flags
If we have PF but no CF this simplifies the IR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 11:01:46 -04:00
Tony Wasserka 527b65648f unittests: Enable logging to stderr when invoking FEXLoader 2023-09-18 16:53:35 +02:00
Tony Wasserka bef64c53f8 unittests: Add test thunk library 2023-09-18 16:53:35 +02:00
Alyssa Rosenzweig 8edcd31404 OpcodeDispatcher: Avoid inverting PF
..if we can fold the invert into the reader.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 10:35:39 -04:00
Ryan Houdek fd1b639ad9 Merge pull request #3115 from lioncash/sqxtun
Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
2023-09-17 14:51:27 -07:00
Ryan Houdek 950a8dbfe7 Merge pull request #3114 from lioncash/ins
Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
2023-09-17 14:40:38 -07:00
Lioncache 26e4d8ad59 Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
If the destination and lower data alias, we can
avoid needing to move into a temporary.
2023-09-17 17:37:36 -04:00
Lioncache d54f590b14 Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
If none of the vectors alias the destination, then we can eliminate
an extra move and usage of a temporary.
2023-09-17 16:56:23 -04:00
Ryan Houdek b3269f20ef Merge pull request #3113 from lioncash/shrn
Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
2023-09-17 13:03:42 -07:00
Lioncache 047646be6d Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
In the event the destination and lower source are the same, then
we don't need to perform any moves.
2023-09-17 15:49:54 -04:00
Ryan Houdek 8168a49d10 Merge pull request #3112 from lioncash/assert
Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
2023-09-17 12:40:47 -07:00
Lioncache 7f2fd4e9a0 Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
Ensures that if our temp vectors change in the future that this is
caught at compile-time rather than runtime.
2023-09-17 15:22:39 -04:00
Lioncache 4ea9f08425 ARMEmitter: Mark index and conversion ops as constexpr
Will be used for assertions. Also makes registers more flexible
for compile-time stuff in general.
2023-09-17 15:19:44 -04:00
Ryan Houdek ffb58761c1 Merge pull request #3111 from lioncash/shift
Arm64/VectorOps: Fix SVE aliasing-path  move in VSShr
2023-09-17 12:15:49 -07:00
Lioncache 8ecdb341e2 Arm64/VectorOps: Fix SVE aliasing-path move in VSShr
Seems like this was a typo from 8d11073, since we'd be moving
into a temporary and then never use it.
2023-09-17 14:43:05 -04:00
Ryan Houdek 0c5c146fcf FEXCore/JitSymbols: Buffer writes to reduce overhead
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.

The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.

To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
  clang benchmark and seeing how frequently it was still writing.
   - 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
  symbols.
   - In most cases 100ms is enough that you won't notice the blip in
     perf.

One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
2023-09-16 17:52:46 -07:00
Ryan Houdek ad8b0c673f Merge pull request #3109 from lioncash/shlx
OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
2023-09-15 18:49:36 -07:00
Ryan Houdek e574cfe681 Merge pull request #3108 from lioncash/mulx
OpcodeDispatcher: Improve output of MULX
2023-09-15 18:09:02 -07:00
Lioncache e9be291cec OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
We can remove some unnecessary moves for the 32-bit cases and
collapse the operations down to a single instruction.
2023-09-15 21:05:50 -04:00
Lioncache d4f87c7db1 OpcodeDispatcher: Improve output of MULX
We can cut down on a few of the generated moves. For
the case where both destinations alias one another,
we can just calculate the high part instead of both of them.
2023-09-15 20:52:02 -04:00
Ryan Houdek 4604c01986 Merge pull request #3107 from lioncash/pext
Arm64/ALUOps: Remove spills in PEXT
2023-09-15 17:40:44 -07:00
Lioncache b0c8ff0ea6 Arm64/ALUOps: Remove spills in PEXT
Reduces the number of emitted instructions for a
corresponding PEXT instruction.

We no longer spill for this IR op.
2023-09-15 19:39:51 -04:00
Ryan Houdek 647629ac23 Merge pull request #3105 from lioncash/rorx
OpcodeDispatcher: Handle RORX corner cases better
2023-09-15 14:55:17 -07:00
Lioncache be90e76422 Arm64/ALUOps: mov in the case of full 32-bit/64-bit BFE
Allows register-renaming mechanisms to be invoked more frequently
2023-09-15 17:38:01 -04:00
Lioncache 4a37ea4819 OpcodeDispatcher: Handle RORX corner cases better
There are a few cases where we were emitting code when we
didn't really need to, or could emit less.
2023-09-15 17:36:36 -04:00
Ryan Houdek 6e08ac65b9 Merge pull request #3106 from lioncash/clwb
HostFeatures: Fix x86 CLWB support check
2023-09-15 14:06:54 -07:00
Lioncache 8705de1893 HostFeatures: Fix x86 CLWB support check
This was clobbering the BMI2 boolean unintentionally.
2023-09-15 16:38:45 -04:00
Alyssa Rosenzweig c8e7c347c3 Merge pull request #3100 from Sonicadvance1/optimize_cmov
OpcodeDispatcher: Optimize cmov
2023-09-15 15:17:51 -04:00
Ryan Houdek 3d0b66407e Merge pull request #3104 from lioncash/vperm2
InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
2023-09-15 11:27:11 -07:00
Lioncache c86b6dc690 InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
Allows viewing the codegen for cases where conditional zeroing is performed.

Also fixes up the vperm2f variants shorthanding one of the registers
to make everything a little more explicit.
2023-09-15 14:10:11 -04:00
Ryan Houdek 773e9465bc Merge pull request #3103 from lioncash/warn
DeadContextStoreElimination: Silence unused function warning
2023-09-15 10:57:40 -07:00
Lioncache e1ed7f43fd DeadContextStoreElimination: Turn LastAccessType into an enum class
Makes the type stricter in terms of implicit conversions.
2023-09-15 13:23:00 -04:00
Ryan Houdek 3eb501aa27 InstCountCI: Update for optimized NZCV and cmov 2023-09-15 10:11:50 -07:00
Ryan Houdek d5b58eebaf OpcodeDispatcher: Optimize cmov
cmov was quite terrible in its implementation. Some things of note:
- NZCV cache would cause store for no reason
- {16,32}-bit would zero extend sources for no reason
- 16-bit would zero extend result for no reason

A bunch of flag testing is still doing a ubfx plus compare against zero
when it could end up being a tst instead, but this is a step in the
right direction and switches over to explicit sized selects.
2023-09-15 10:09:37 -07:00
Ryan Houdek 6dbbd9ecfc OpcodeDispatcher: Duplicate SelectCC but with Explicit result size
This is a temporary measure as we are moving Select operations over to
explicit sizes. Once we remove all uses of SelectCC then it will get
removed.
2023-09-15 10:09:37 -07:00
Ryan Houdek 759cc0025a OpcodeDispatcher: Add a dirty flag for tracking NZCV status
Cached NZCV reads don't need to be written back at the end of the block.
This will remove one instruction from the end of some blocks.
2023-09-15 10:09:37 -07:00
Lioncache d05f890147 DeadContextStoreElimination: Silence unused function warning 2023-09-15 13:08:32 -04:00
Alyssa Rosenzweig 9152fb030e Merge pull request #3102 from alyssarosenzweig/inline-xor
Inline constant with PF calculation
2023-09-15 12:48:33 -04:00
Alyssa Rosenzweig 2385c275ac InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:33:53 -04:00
Alyssa Rosenzweig d29b8bab36 OpcodeDispatcher: Inline constant in PF calculation
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:33:53 -04:00
Ryan Houdek b6922dff57 Merge pull request #3101 from alyssarosenzweig/opt/dec
Optimize out carry invert for DEC
2023-09-15 09:30:08 -07:00
Alyssa Rosenzweig c560a88de4 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:10:57 -04:00
Alyssa Rosenzweig fc02f38435 IR: Only invert CF for NZCV if needed
If we are going to throw away the updated value of CF anyway there is no point
wasting an instruction to invert CF. Add an IR toggle for that so the arm64 JIT
can make better choices.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:08:22 -04:00
Ryan Houdek d5782567e8 Merge pull request #3077 from Sonicadvance1/x86_shifted
FEXCore: Implements support for shifted bitwise ops
2023-09-15 08:09:35 -07:00
Ryan Houdek 060433621a Merge pull request #3097 from Sonicadvance1/disable_enhanced_tso
FEXCore: Disable Enhanced REP MOVSB if Atomic TSO is enabled
2023-09-15 08:08:36 -07:00
Ryan Houdek 9866e238d5 Merge pull request #3080 from Sonicadvance1/defer_softfloat
FEXCore: Defer setting x87 softflow rounding mode until use
2023-09-15 08:08:04 -07:00
Mai f5c4e28696 Merge pull request #3098 from Sonicadvance1/optimize_vectors_sve
Arm64: Optimize wide shifts slightly for 64-bit OpSize
2023-09-15 05:21:20 -04:00
Mai 96bbd01ad6 Merge pull request #3096 from Sonicadvance1/optimal_crc
OpcodeDispatcher: Optimize CRC32
2023-09-15 05:19:05 -04:00
Mai f84a264b0e Merge pull request #3095 from Sonicadvance1/bswap
OpcodeDispatcher: Optimize 16-bit MOVBE
2023-09-15 05:18:26 -04:00
Mai a8c17201b5 Merge pull request #3099 from Sonicadvance1/explicit_but_implicit_select
IR: Changes Select operation to not have implicit sizes
2023-09-15 05:17:47 -04:00
Ryan Houdek d81d89c4fb IR: Changes Select operation to not have implicit sizes
Changes the helper which all the source uses to still calculate the size
implicitly. This is going to take a while to convert all implicit uses
over to the explicit operation.

Get us started by at least having the IR operation itself be explicit.
2023-09-14 20:48:16 -07:00
Ryan Houdek 92212c48f1 IR: Fixes parsing of default arguments with colons
We need to split on the first colon, not every colon in the arguments.

This will be used in the next changes.
2023-09-14 20:37:45 -07:00
Ryan Houdek 42a24bbbd1 InstCountCI: Update CI for optimized wide shifts 2023-09-14 19:53:18 -07:00
Ryan Houdek 021c99e233 Arm64: Optimize wide shifts slightly for 64-bit OpSize
Wide shifts under SVE use 64-bit source elements. If a smaller element
overlaps the 64-bit shift element then it uses that shift
eg:
- Src1[15:0] >> Shift[63:0]
- Src1[31:16] >> Shift[63:0]
- Src1[47:32] >> Shift[63:0]
- Src1[63:48] >> Shift[63:0]
- After this point it will switch to the next 64-bit shift element
- Src1[79:64] >> Shift[127:64]
- Src1[95:80] >> Shift[127:64]
- Src1[111:96] >> Shift[127:64]
- Src1[127:112] >> Shift[127:64]

As seen, we can skip the duplication of the scalar element if the OpSize
is 64-bit, this makes MMX emulation slightly more optimal here.
This also means that a few instructions that weren't claimed to be
optimal actually are since they need the duplication operation (which
vixl always labels as a mov).
2023-09-14 19:35:56 -07:00
Ryan Houdek f730339365 OpcodeDispatcher: Reorder vector loads in shifts
This affects codegen due to RA quirks. This ensures that the wide shifts
don't have to generate a movprfx.
2023-09-14 19:34:00 -07:00
Ryan Houdek 40a4eb90af FEXCore: Disable Enhanced REP MOVSB if TSO is enabled
Hades and the vcruntime hits this very hard in memmove.

`86.56%  [JIT] tid 458574        [.] JIT_0x18000c375_0x7fffc94790c8`

```asm
   0x00007fffc94790f8:  ldaprb  w3, [x2]
   0x00007fffc94790fc:  stlrb   w3, [x1]
   0x00007fffc9479100:  add     x1, x1, #0x1
   0x00007fffc9479104:  add     x2, x2, #0x1
   0x00007fffc9479108:  sub     x0, x0, #0x1
   0x00007fffc947910c:  cbnz    x0, 0x7fffc94790f8
```

This performance is terrible because Cortex's LRCPC performance is bottom-tier.
Work around the performance issue by forcing things to do larger moves with vector moves instead.
2023-09-14 17:23:02 -07:00
Ryan Houdek 31ad26202e InstCountCI: Update for Optimized CRC32 2023-09-14 16:23:33 -07:00
Ryan Houdek a3115d4699 OpcodeDispatcher: Optimize CRC32
The only version of this instruction that was generating optimal code
was the one with 64-bit destination and source.

Optimizes the rest of the operating sizes so that they are all optimal
at one instruction translations
2023-09-14 16:21:42 -07:00
Ryan Houdek c3ead80927 InstCountCI: Update for optimized 16-bit movbe 2023-09-14 16:05:38 -07:00
Ryan Houdek 80cda1bb18 OpcodeDispatcher: Optimize 16-bit MOVBE
16-bit MOVBE is a bit of a special case where it loads 16-bits in to the
bottom of the GPR without clearing the upper bits of the register.
Which means 32-bits or 64-bits depending on operating mode.

Arm64 doesn't support a 16-bit bswap so it needs to operate at 32-bits
instead. We then can insert the resulting bits of the 32-bit rev with a
bfxil in to the lower bits of the resulting destination register.

This allows 16-bit movbe to be optimal now.
2023-09-14 16:05:16 -07:00
Ryan Houdek 90ddee5f8d IR: Implements support for arm64 bfxil
This is useful for extracting a width from a register and inserting in
to the lower bits of a destination.
2023-09-14 15:52:55 -07:00
Ryan Houdek 6fdf2f963b Merge pull request #3082 from Sonicadvance1/minor_storeregsra_opt
FEXCore: Minor optimization to StoreRegisterSRA
2023-09-14 14:54:39 -07:00
Ryan Houdek e1eb151051 Merge pull request #3094 from neobrain/refactor_reorder_ci
CI: Run tests with <30s runtime first
2023-09-14 13:56:12 -07:00
Tony Wasserka 3f8bf01f75 CI: Run tests with <30s runtime first 2023-09-14 20:46:50 +02:00
Mai 92824f5e4d Merge pull request #3093 from Sonicadvance1/optimize_blendp
OpcodeDispatcher: Optimize blendp{s,d}
2023-09-14 00:33:00 -04:00
Mai 213d3c4e2b Merge pull request #3091 from Sonicadvance1/optimize_pinsr
OpcodeDispatcher: Optimize pins{b,w,d,q}
2023-09-14 00:30:56 -04:00
Mai d4c6749d2a Merge pull request #3090 from Sonicadvance1/optimize_pextr
OpcodeDispatcher: Optimize pextr{b,w}
2023-09-14 00:30:43 -04:00
Mai 1804b007ec Merge pull request #3092 from Sonicadvance1/instcountci_compile_log
InstCountCI: Add log before compiling instruction
2023-09-13 23:14:42 -04:00
Mai 655cee070d Merge pull request #3089 from Sonicadvance1/optimize_pshufd
OpcodeDispatcher: Optimize shufpd
2023-09-13 23:12:10 -04:00
Ryan Houdek 28309a1cc5 InstCountCI: Update for blend 2023-09-13 20:08:03 -07:00
Ryan Houdek 29f824cf7a OpcodeDispatcher: Optimize blendp{s,d}
Optimal blendps is worst case 2 instructions.
FEX's RA doesn't quite get there since it can't see through multiple
instructions with SRA destinations. That'll be fixed in the future.

Optimal blendps is always one instruction, one is a no-op.
We always hit this.
2023-09-13 20:06:39 -07:00
Ryan Houdek 2bafa2c26f InstCountCI: Update for optimized pins{b,w,d,q} 2023-09-13 19:53:05 -07:00
Ryan Houdek 5e7d793a6a OpcodeDispatcher: Optimize pins{b,w,d,q}
Inserting from a GPR and memory can both be optimized. These are now
optimal

Needs #3088 merged first.
2023-09-13 19:53:05 -07:00
Ryan Houdek 3e40713ccc InstCountCI: Update for optimized pextr{b,w} 2023-09-13 19:51:56 -07:00
Ryan Houdek 33a2fbb896 OpcodeDispatcher: Optimize pextr{b,w}
Cleans up the code which had special cased some 32-bit optimization
which is unnecessary now that both 8-bit and 16-bit are also optimized.

When FEX does a VExtractToGPR, the result is zero extended to the full
GPR register size. This means we don't need to do a zero extend when
storing to a guest GPR.

Makes pextr{b,w} optimal now.

Needs #3088 merged first.
2023-09-13 19:51:56 -07:00
Ryan Houdek 853ded7df7 InstcountCI: Update for optimized shufpd 2023-09-13 19:50:15 -07:00
Ryan Houdek 67914157cb OpcodeDispatcher: Optimize shufpd
This one is very satisfying since there are only four variants and each
one of them converts to a single instruction.

Needs #3088 merged first
2023-09-13 19:50:15 -07:00
Mai 750d90939d Merge pull request #3088 from Sonicadvance1/instcountci_missing_secondary_opsize
InstCountCI: Adds missing instructions from Secondary OpSize tables
2023-09-13 22:49:17 -04:00
Mai 31d828390f Merge pull request #3087 from Sonicadvance1/tbl2_implementation
OpcodeDispatcher: Implement shufps with VTBL2 in worst case
2023-09-13 22:48:53 -04:00
Ryan Houdek 6c2f8ab085 InstCountCI: Add log before compiling instruction
If CI faults out due to a bug then we would have no log as to which
instruction caused the issue.

I find myself adding this each time an assert fires to see what
instruction it was working on. Just add it directly.
2023-09-13 14:33:33 -07:00
Ryan Houdek 2aea401189 InstCountCI: Adds missing instructions from Secondary OpSize tables
I managed to miss a whole section of instructions from the secondary
opsize tables. This resulted in four instructions missing from the
database.

Adds cmppd, pinsrw, pextrw, and shufpd which are all non-optimal
instruction implementations.
2023-09-13 11:48:50 -07:00
Ryan Houdek c008671509 unittests/asm: Add test with inverted sources
To ensure this is tested with non sequential source registers.
2023-09-13 11:31:20 -07:00
Ryan Houdek 5903be156c InstCountCI: Update for shufps tbl opt 2023-09-13 11:31:20 -07:00
Ryan Houdek db5056f275 OpcodeDispatcher: Implement shufps with VTBL2 in worst case
In the case that source registers are sequential then this turns in to a
load of the vector constant (2 instructions) and the single tbl
instruction.

If the registers aren't sequential then the tbl turns in to 2 moves and
then the single tbl, which with zero-cycle rename isn't too bad.

Since this is a worst case option this is significantly better than the
previous implementation doing a bunch of inserts which was always 9
instructions.
We should still strive to implement faster versions without the use of
TBL2 if possible but this makes it less of a concern.
2023-09-13 11:31:20 -07:00
Ryan Houdek e9d96ce538 IR: Implements support for VTBL2
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.

The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
2023-09-13 11:31:20 -07:00
Ryan Houdek 444d4c082d Int: Fixes typo in LoadNamedVectorIndexedConstant
Surprising this didn't break anything before this.
2023-09-13 11:31:20 -07:00
Ryan Houdek cfe620ab15 Merge pull request #3085 from Sonicadvance1/optimize_shufps
OpcodeDispatcher: Optimize a bunch of shufps variants
2023-09-12 21:53:33 -07:00
Ryan Houdek ea8d63350a InstCountCI: Updates for optimized shufps 2023-09-12 19:58:07 -07:00
Ryan Houdek e37cef8283 unittests: Implement shufps optimization test
Tests all current forms of shufps optimizations.
2023-09-12 19:58:07 -07:00
Ryan Houdek 3f1979286f OpcodeDispatcher: Optimize a bunch of shufps variants
Hits a whole bunch of common cases, most of which then emit optimal code
generation.
Two cases that use VInsElement hit the RA quirk where the SRA
destination is dead but RA doesn't see it, so it ends up doing a couple
moves. If RA gets fixed then those two moves will go away.

There are definitely still cases that we could emit more optimal code.
Additionally we could implement a TBL2 IR operation to do a LUT approach
for ones we don't cover.

Problem with implementing a TBL2 ir operation is that we have no way to
ensure registers are sequential so we would need to always do moves
```asm
ldr v2, <LUT Table>
mov v0, v16
mov v1, v18
tbl v16.16b, { v0.16b, v1.16b }, v2.16b
```

Which to be fair isn't terrible, and if we're lucky that the guest uses
sequential registers we can naturally get the more optimal code path.
Ideally our RA could push some operations in to sequential registers but
that's not possible currently.

I'll do a follow-up PR that implements TBL2.
2023-09-12 19:58:07 -07:00
Ryan Houdek d5c3036bc2 JITx86: Fixes VREV64 with 32-bit element size.
This has been incorrect since it has been implemented.
Noticed when implementing optimizations.
2023-09-12 19:23:59 -07:00
Mai ebdca02218 Merge pull request #3084 from Sonicadvance1/optimize_bswap
OpcodeDispatcher: Optimize 32-bit bswap
2023-09-12 20:09:00 -04:00
Mai dda5861bdd Merge pull request #3081 from Sonicadvance1/fix_waitpid
Tools: Fixes usage of waitpid in the face of EINTR
2023-09-12 19:35:05 -04:00
Mai f7e652b616 Merge pull request #3083 from Sonicadvance1/optimize_nop_move
OpcodeDispatcher: Optimize NOP vector move
2023-09-12 19:34:36 -04:00
Ryan Houdek 65bc159ff1 InstCountCI: Update for bswap optimization 2023-09-12 16:19:55 -07:00
Ryan Houdek c362d3a9d8 OpcodeDispatcher: Optimize 32-bit bswap
Removes a redundant move, making it optimal now.
2023-09-12 16:19:10 -07:00
Ryan Houdek 8a44be0c30 InstCountCI: Update for NOP vector moves
Adds a couple of instructions that get tested in this code path.
2023-09-12 16:11:43 -07:00
Ryan Houdek 304dba5f20 OpcodeDispatcher: Optimize NOP vector move
Move instruction to itself here is a nop.
Need to be careful about AVX operations which use a different handler
since those might actually zero the upper bits on 128-bit move
2023-09-12 16:10:39 -07:00
Ryan Houdek b2e61d2deb InstCountCI: Update for minor storeregistersra opt 2023-09-12 14:38:46 -07:00
Ryan Houdek e6c0bebee9 FEXCore: Minor optimization to StoreRegisterSRA
{Load,Store}RegisterSRA always loads or stores GPRSize. 8-bit and 16-bit
are vestigial and all OpcodeDispatcher usage will load the full GPR size
(32-bit or 64-bit) and then extract or insert as necessary.

This cleans up a few bits of codegen in InstCountCI.
2023-09-12 14:36:08 -07:00
Ryan Houdek aa017116b3 Tools: Fixes usage of waitpid in the face of EINTR
waitpid can return early if interrupted due to EINTR.
Loop on this case and try again.
2023-09-12 12:41:43 -07:00
Ryan Houdek 97a6184e53 InstCountCI: Update for FCW optimization 2023-09-12 05:21:06 -07:00
Ryan Houdek 76bd81af15 FEXCore: Defer setting x87 softflow rounding mode until use
Currently FEX will always jump out of the JIT any time FCW was getting
written to, ensuring that the softfloat state is setup to rounding at
the time of FCW getting written.
This has the unintended side-effect that even in "x87 reduced precision"
mode we were jumping out of the JIT.
This hit a real world use case of an installer reloading FCW after every
x87 operation and generating a block with 2297 instructions.

Instead when jumping out of the JIT for handling x87 operations, load
FCW and pass it as the first argument of the handler. Setting the
softfloat state at that point.

This helps the installer's hottest block by cutting it down to 1477
instructions. 64.3% of the original size. The code block is still
burning 90% of the CPU time of the installer but the performance is
significantly better while it is doing its decompression.

In order to optimize this installer's block of code more then we will
likely need to optimize out x87 stack usage.
2023-09-12 05:21:06 -07:00
Mai 90f7937146 Merge pull request #3079 from Sonicadvance1/recover_two_temps
Arm64: Recover two unused vector vector temporary registers
2023-09-11 22:06:03 -04:00
Mai 98f148766d Merge pull request #3078 from Sonicadvance1/detect_flagm
HostFeatures: Detect FlagM/2
2023-09-11 20:57:43 -04:00
Ryan Houdek 9c44e295fa InstCountCI: Update for recovering two vector temps
All the changes are RA changes and spilling/filling taking another
instruction.
2023-09-11 16:50:52 -07:00
Ryan Houdek b5a1d323c2 Arm64: Recover two unused vector vector temporary registers
This leaves us with two temporary vectors that the JIT can use.
As of last month we stopped using v2 and v3 as temporaries and these can
now be given back to the JIT.

Ensures that the registers are still sequentially ordered and adds
support for spilling the FPR counts that are aligned by 2 instead of 4.
Adds a couple of instructions to filling and spilling but isn't that big
of an issue.

InstcountCI has some ridiculously large changes just because RA is
starting at a new register number.
2023-09-11 16:48:25 -07:00
Ryan Houdek b453439968 HostFeatures: Detect FlagM/2
Currently unused but at least detect the feature so that our Arm64 JIT
can use it in the future.
2023-09-11 16:41:30 -07:00
Ryan Houdek 863331b117 FEXCore: Implements support for shifted bitwise ops
This wasn't implemented initially for the interpreter and x86 JIT.

This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.

The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
2023-09-11 13:17:35 -07:00
Mai 48521a4416 Merge pull request #3075 from Sonicadvance1/optimize_bt_ops
OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
2023-09-11 16:05:33 -04:00
Mai 6fe643d270 Merge pull request #3076 from Sonicadvance1/enable_enhanced_rep_movs
CPUID: Enabled Enhanced REP MOVSB/STOSB
2023-09-11 15:35:56 -04:00
Mai fbc4bda7a6 Merge pull request #3074 from Sonicadvance1/hwcap2_fsgsbase
ELFCodeLoader: Expose FSGSBase in getauxval HWCAP2
2023-09-11 15:35:26 -04:00
Mai 6d9b52452e Merge pull request #3072 from Sonicadvance1/crc32_is_fixed_size
IR: Changes crc32 operation to always return a 32-bit result.
2023-09-11 15:34:37 -04:00
Mai 950007c815 Merge pull request #3071 from Sonicadvance1/update_rcl_opsize
OpcodeDispatcher: Update 32/64-bit RCL for operating size
2023-09-11 15:34:06 -04:00
Mai d029394c27 Merge pull request #3070 from Sonicadvance1/update_rcr_opsize
OpcodeDispatcher: Update 32/64-bit RCR for operating size
2023-09-11 15:33:34 -04:00
Mai 879fcdc6fe Merge pull request #3069 from Sonicadvance1/fix_redundant_load_rclse
IR:RCLSE: Partially reenables the RCLSE pass
2023-09-11 15:32:55 -04:00
Ryan Houdek 2f77982b54 CPUID: Enabled Enhanced REP MOVSB/STOSB
Missed with #2490.
This changes behaviour of glibc's memmove slightly, seems to recover a
bit of performance on Half-Life 2's title screen.
2023-09-10 20:49:08 -07:00
Ryan Houdek e3a00fb2fb InstCountCI: Update for BT minor opt 2023-09-10 20:23:08 -07:00
Ryan Houdek 4feb059f51 OpcodeDispatcher: Optimize the case of all flags invalidated
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
2023-09-10 20:16:29 -07:00
Ryan Houdek 3d1bbe505d OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
These instructions set all the flags to undefined and moves the
resulting bit in to CF. No need to calculate the deferred flags when
we are about to write over them.
2023-09-10 20:16:29 -07:00
Ryan Houdek b2a42b6c61 ELFCodeLoader: Expose FSGSBase in getauxval HWCAP2
We have supported this since #163 but we haven't been exposing the
feature in hwcap2.

We have exposed it in CPUID this entire time, just not in hwcap2.
2023-09-10 17:00:21 -07:00
Ryan Houdek 315d1855de IR: Changes crc32 operation to always return a 32-bit result.
CRC32 is always a 32-bit sized operation even with a 64-bit source
value.
This doesn't change any InstCountCI results.
2023-09-09 10:02:30 -07:00
Ryan Houdek 93246878e2 InstCountCI: Update for rcl explicit size change 2023-09-09 09:40:36 -07:00
Ryan Houdek 6c62691af0 OpcodeDispatcher: Update 32/64-bit RCL for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:40:12 -07:00
Ryan Houdek ee5aed51d8 InstCountCI: Update for rcr expliti size change 2023-09-09 09:35:13 -07:00
Ryan Houdek 47f50a7008 OpcodeDispatcher: Update 32/64-bit RCR for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:33:27 -07:00
Ryan Houdek be07254935 Merge pull request #3067 from neobrain/refactor_thunks
Thunks: Minor restructuring and small cleanups
2023-09-07 20:16:17 -07:00
Ryan Houdek 636f8aa4a7 Arm64: Fix undefined behaviour in Push operation
Arm64 store with writeback when source register is the same register as
the address is undefined behaviour.
Depending on hardware details this can do a whole bunch of things.

This situation happens when the x86 code does `push rsp` which is quite
common for applications to do. We would then convert this to a `str x8, [x8, #-8]!`
Which results in undefined behaviour.

Now that redundant loads are optimized this showed up as an issue. Adds
a unit test to ensure we don't hit this again.
2023-09-07 17:38:39 -07:00
Ryan Houdek 22ca46a227 Arm64: Fixes SVE V{S,U}MulH
When the destination overlaps one of the sources we must be careful to
follow a movprfx rule.
```
The destination register must not refer to architectural register state
referenced by any other source operand register of this instruction.
```

We ended up in a situation in the vpmulh{u,}w AVX tests where zm was
overlapping the destination which violated that rule. This also
generated invalid code for this instruction.
```
[INFO] movprfx z6, z4
[INFO] umulh z6.h, p6/m, z6.h, z6.h
```

As seen, we were overwriting one of the sources because the destination
overlapped it. Now instead check if each individual overlap so invalid
code isn't generated.

InstCountCI results aren't affected since this only happens in
situations with multiple instructions.
2023-09-07 16:47:08 -07:00
Ryan Houdek 7b80427de0 OpcodeDispatcher: Remove BLENDV "optimization"
Now that the RCLSE pass finally optimizes redundant loads again this
optimization that lives in the OpcodeDispatcher can be removed.

With InstCountCI reran, the pblendvb results don't change at all, as
expected.
2023-09-07 16:00:56 -07:00
Ryan Houdek b753b9ffa2 InstCountCI: Updates for RCLSE fix
Adds two `packsswb` tests to ensure redundant sources are getting
optimized as expected.
2023-09-07 15:58:43 -07:00
Ryan Houdek c62b5a3103 IR:RCLSE: Partially reenables the RCLSE pass
This is taking steps to start fixing RCLSE which was started by #2700.
Same situation as that PR, since #2170 when we converted
{Load,Store}Context in to {Load,Store}Register we broke this pass
entirely. It hasn't been doing anything for redundant GPRs and FPRs
since at least November of last year.

Technically it was potentially still optimizing redundant MMX
accesses, but it is so broken that it doesn't matter.

Instead of going all in like #2700 did, tear down the pass and start
again. We are now /only/ optimizing redundant context/register loads.
This fixes an issue that comes up commonly where the same register used
as sources was getting loaded twice, causing redundant moves.

`packsswb xmm0, xmm0` for example was generating a four instruction
sequence instead of three instructions because we weren't eliminating
the redundant load.

Going to take reimplementing all the optimizations that this pass does
in steps. This way we can track any regression in the independent steps
unlike what happened in #2700.

Confirmed that Proton/Sonic Mania still works after this.
2023-09-07 15:50:27 -07:00
Ryan Houdek 3c729bcacb IR/RCLSE: Removes unused CalculateControlFlowInfo
This is unused and this only optimizes inside of a block.
2023-09-07 15:49:29 -07:00
Tony Wasserka 677b77f1bb Thunks: Simplify PackedArguments invocation code 2023-09-07 13:56:53 +02:00
Tony Wasserka 024fb268c0 unittests/ThunkLibs: Move utility code to a dedicated header 2023-09-07 13:56:53 +02:00
Tony Wasserka ab8bc052a0 Thunks/gen: Split interface parsing and code emission into separate files 2023-09-07 13:56:53 +02:00
Tony Wasserka 971821460c Thunks/gen: Move diagnostic marshaling helper to a dedicated header 2023-09-07 13:56:53 +02:00
Ryan Houdek 615ab8d80c Merge pull request #3066 from neobrain/fix_procfsint_regression
FileManagement: Fix inverted boolean check for procfs/interpreter support
2023-09-06 17:37:12 -07:00
Tony Wasserka 83c74e86c8 FileManagement: Fix inverted boolean check for procfs/interpreter support 2023-09-06 17:00:58 +02:00
Mai 9ff5544d55 Merge pull request #3065 from Sonicadvance1/fix_docs_location
Scripts: Update generate_doc_outline for moved FEXCore
2023-09-06 10:54:20 -04:00
Ryan Houdek 2b9265d9fe Scripts: Update generate_doc_outline for moved FEXCore
Otherwise all of the FEXCore docs get deleted.
2023-09-05 23:23:27 -07:00
741 changed files with 117542 additions and 46731 deletions

No files matched your search

+50 -50
View File
@@ -17,11 +17,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
build_plus_test:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64], [self-hosted, ARMv8.0], [self-hosted, ARMv8.2], [self-hosted, ARMv8.4]]
arch: [[self-hosted, ARMv8.0], [self-hosted, ARMv8.2], [self-hosted, ARMv8.4]]
fail-fast: false
steps:
@@ -65,7 +65,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -78,18 +78,6 @@ jobs:
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -102,30 +90,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gvisor tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gvisor_tests
- name: GVisor Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GVisor.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -150,17 +114,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC32.log || true
- name: Struct verifier tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target struct_verifier
- name: Struct verifier Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -244,6 +197,53 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkResults.log || true
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gvisor tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gvisor_tests
- name: GVisor Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GVisor.log || true
- name: Struct verifier tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target struct_verifier
- name: Struct verifier Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+27 -28
View File
@@ -24,12 +24,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
glibc_fault_test:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
# Run on an x86 device and any ARM runner.
arch: [[self-hosted, x64], [self-hosted, ARM64]]
arch: [[self-hosted, ARM64]]
fail-fast: false
steps:
@@ -73,7 +72,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -86,18 +85,6 @@ jobs:
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -110,18 +97,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -179,6 +154,30 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+107
View File
@@ -0,0 +1,107 @@
name: Hostrunner tests
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
# Customize the CMake build type here (Release, Debug, RelWithDebInfo, etc.)
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
hostrunner_tests:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+41 -3
View File
@@ -16,11 +16,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
instcountci_tests:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64]]
arch: [[self-hosted, x64], [self-hosted, ARM64]]
fail-fast: false
steps:
@@ -56,6 +56,16 @@ jobs:
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Set vixl_sim x86
if: matrix.arch[1] == 'x64'
run: |
echo "VIXL_SIM_ENABLED=True" >> $GITHUB_ENV
- name: Set vixl_sim Arm64
if: matrix.arch[1] == 'ARM64'
run: |
echo "VIXL_SIM_ENABLED=False" >> $GITHUB_ENV
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
@@ -64,7 +74,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=False -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -86,6 +96,25 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_InstCountCI.log || true
- name: Update local repo instcount
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target instcountci_update_tests
- name: Get instcountCI diff
if: ${{ always() }}
shell: bash
working-directory: ${{github.workspace}}/
run: git diff --output=${{runner.workspace}}/build/InstCountCI.diff
- name: Check if InstCountCI Diff exists
if: ${{ always() }}
shell: bash
working-directory: ${{github.workspace}}/
# Check if the file is empty
run: sh -c "! test -s ${{runner.workspace}}/build/InstCountCI.diff"
- name: Truncate test results
if: ${{ always() }}
shell: bash
@@ -107,3 +136,12 @@ jobs:
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
- name: Upload results InstCountCI
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}-instcountci
path: ${{runner.workspace}}/build/InstCountCI.diff
retention-days: 3
+3 -3
View File
@@ -13,11 +13,11 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
mingw_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64, mingw], [self-hosted, ARM64, mingw]]
arch: [[self-hosted, ARM64, mingw]]
fail-fast: false
steps:
@@ -74,7 +74,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=False -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+1 -1
View File
@@ -16,7 +16,7 @@ env:
FEX_ENABLEAVX: 1
jobs:
build:
vixl_simulator:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
+5
View File
@@ -0,0 +1,5 @@
{
"ThunksDB": {
"fex_thunk_test": 1
}
}
+8 -21
View File
@@ -25,7 +25,6 @@ option(ENABLE_JEMALLOC_GLIBC_ALLOC "Enables jemalloc glibc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
option(ENABLE_VIXL_SIMULATOR "Forces the FEX JIT to use the VIXL simulator" FALSE)
@@ -97,11 +96,6 @@ if (ENABLE_GDB_SYMBOLS)
endif()
if (ENABLE_INTERPRETER)
message(STATUS "Interpreter enabled")
add_definitions(-DINTERPRETER_ENABLED=1)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -118,14 +112,6 @@ else()
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
if (NOT ENABLE_X86_HOST_DEBUG)
message(FATAL_ERROR
" Be warned: FEX isn't optimized for x86_64 hosts!\n"
" Support for x86_64 hosts is only for debugging and convenience!\n"
" Don't expect amazing performance or optimal code generation!\n"
" Pass -DENABLE_X86_HOST_DEBUG=True to bypass this message!")
endif()
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
@@ -307,10 +293,11 @@ if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
endif()
endif()
set(FEX_TUNE_COMPILE_FLAGS)
if (NOT TUNE_ARCH STREQUAL "generic")
check_cxx_compiler_flag("-march=${TUNE_ARCH}" COMPILER_SUPPORTS_ARCH_TYPE)
if(COMPILER_SUPPORTS_ARCH_TYPE)
add_compile_options("-march=${TUNE_ARCH}")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=${TUNE_ARCH}")
else()
message(FATAL_ERROR "Trying to compile arch type '${TUNE_ARCH}' but the compiler doesn't support this")
endif()
@@ -323,7 +310,7 @@ if (TUNE_CPU STREQUAL "native")
# Clang can not currently check for native Apple M1 type in hypervisor. Currently disabled
check_cxx_compiler_flag("-mcpu=native" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=native")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=native")
endif()
else()
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
@@ -337,19 +324,19 @@ if (TUNE_CPU STREQUAL "native")
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=${AARCH64_CPU}")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${AARCH64_CPU}")
endif()
endif()
else()
check_cxx_compiler_flag("-march=native" COMPILER_SUPPORTS_MARCH_NATIVE)
if(COMPILER_SUPPORTS_MARCH_NATIVE)
add_compile_options("-march=native")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=native")
endif()
endif()
else()
check_cxx_compiler_flag("-mcpu=${TUNE_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=${TUNE_CPU}")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${TUNE_CPU}")
else()
message(FATAL_ERROR "Trying to compile cpu type '${TUNE_CPU}' but the compiler doesn't support this")
endif()
@@ -467,10 +454,10 @@ if (BUILD_THUNKS)
CMAKE_ARGS
"-DBITNESS=64"
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DBUILD_FEX_LINUX_TESTS=${BUILD_FEX_LINUX_TESTS}"
"-DENABLE_CLANG_THUNKS=${ENABLE_CLANG_THUNKS}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_64_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
@@ -485,10 +472,10 @@ if (BUILD_THUNKS)
CMAKE_ARGS
"-DBITNESS=32"
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DBUILD_FEX_LINUX_TESTS=${BUILD_FEX_LINUX_TESTS}"
"-DENABLE_CLANG_THUNKS=${ENABLE_CLANG_THUNKS}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_32_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"Env": "STEAM_GAME_LAUNCH_SHELL=@CMAKE_INSTALL_PREFIX@/bin/FEXBash"
}
}
+6
View File
@@ -144,6 +144,12 @@
"@PREFIX_LIB@/libasound.so.2.0.0"
]
},
"fex_thunk_test": {
"Library": "libfex_thunk_test-guest.so",
"Overlay": [
"@PREFIX_LIB@/libfex_thunk_test.so"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
-13
View File
@@ -1,13 +0,0 @@
DO WHAT THE FUCK YOU WANT TO PUBLIC LICENSE
Version 2, December 2004
Copyright (C) 2018 Ryan Houdek <Sonicadvance1@gmail.com>
Everyone is permitted to copy and distribute verbatim or modified
copies of this license document, and changing it is allowed as long
as the name is changed.
DO WHAT THE FUCK YOU WANT TO PUBLIC LICENSE
TERMS AND CONDITIONS FOR COPYING, DISTRIBUTION AND MODIFICATION
0. You just DO WHAT THE FUCK YOU WANT TO.
+1 -1
+16 -9
View File
@@ -13,15 +13,6 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
set(_M_ARM_64 1)
endif()
if (ENABLE_VIXL_SIMULATOR)
# If the vixl simulator is enabled then we are using the ARM64 JIT
option(ENABLE_JIT_X86_64 "Enable the x86_64 JIT" FALSE)
option(ENABLE_JIT_ARM64 "Enable the ARM64 JIT" TRUE)
else()
option(ENABLE_JIT_X86_64 "Enable the x86_64 JIT" ${_M_X86_64})
option(ENABLE_JIT_ARM64 "Enable the ARM64 JIT" ${_M_ARM_64})
endif()
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
set(CMAKE_POSITION_INDEPENDENT_CODE ON)
@@ -33,6 +24,22 @@ set(CMAKE_INCLUDE_CURRENT_DIR ON)
include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
include(CheckCXXSourceCompiles)
set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
check_cxx_source_compiles(
"
__attribute__((preserve_all))
void Testy() {
}
int main() {
return 0;
}"
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
if (HAS_CLANG_PRESERVE_ALL)
message(STATUS "Has clang::preserve_all")
endif ()
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
+73 -10
View File
@@ -46,11 +46,15 @@ class OpDefinition:
NumElements: str
OpClass: str
HasSideEffects: bool
ImplicitFlagClobber: bool
RAOverride: int
SwitchGen: bool
ArgPrinter: bool
SSAArgNum: int
NonSSAArgNum: int
DynamicDispatch: bool
JITDispatch: bool
JITDispatchOverride: str
Arguments: list
EmitValidation: list
Desc: list
@@ -64,11 +68,15 @@ class OpDefinition:
self.OpClass = None
self.OpSize = 0
self.HasSideEffects = False
self.ImplicitFlagClobber = False
self.RAOverride = -1
self.SwitchGen = True
self.ArgPrinter = True
self.SSAArgNum = 0
self.NonSSAArgNum = 0
self.DynamicDispatch = False
self.JITDispatch = True
self.JITDispatchOverride = None
self.Arguments = []
self.EmitValidation = []
self.Desc = []
@@ -144,7 +152,7 @@ def parse_ops(ops):
Argument = Argument.strip()
OpArg = OpArgument()
Split = Argument.split(":")
Split = Argument.split(":", 1)
if len(Split) != 2:
ExitError("Error parsing argument. Missing Type and name colon split")
@@ -213,6 +221,9 @@ def parse_ops(ops):
if "HasSideEffects" in op_val:
OpDef.HasSideEffects = bool(op_val["HasSideEffects"])
if "ImplicitFlagClobber" in op_val:
OpDef.ImplicitFlagClobber = bool(op_val["ImplicitFlagClobber"])
if "ArgPrinter" in op_val:
OpDef.ArgPrinter = bool(op_val["ArgPrinter"])
@@ -228,6 +239,15 @@ def parse_ops(ops):
if "Desc" in op_val:
OpDef.Desc = op_val["Desc"]
if "DynamicDispatch" in op_val:
OpDef.DynamicDispatch = bool(op_val["DynamicDispatch"])
if "JITDispatch" in op_val:
OpDef.JITDispatch = bool(op_val["JITDispatch"])
if "JITDispatchOverride" in op_val:
OpDef.JITDispatchOverride = op_val["JITDispatchOverride"]
# Do some fixups of the data here
if len(OpDef.EmitValidation) != 0:
for i in range(len(OpDef.EmitValidation)):
@@ -357,6 +377,7 @@ def print_ir_sizes():
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetRAArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool HasSideEffects(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool ImplicitFlagClobber(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool GetHasDest(IROps Op);\n")
output_file.write("#undef IROP_SIZES\n")
@@ -450,15 +471,17 @@ def print_ir_getraargs():
def print_ir_hassideeffects():
output_file.write("#ifdef IROP_HASSIDEEFFECTS_IMPL\n")
output_file.write("constexpr std::array<uint8_t, OP_LAST + 1> SideEffects = {\n")
for op in IROps:
output_file.write("\t{},\n".format(("true" if op.HasSideEffects else "false")))
for array, prop in [("SideEffects", "HasSideEffects"),
("ImplicitFlagClobbers", "ImplicitFlagClobber")]:
output_file.write(f"constexpr std::array<uint8_t, OP_LAST + 1> {array} = {{\n")
for op in IROps:
output_file.write("\t{},\n".format(("true" if getattr(op, prop) else "false")))
output_file.write("};\n\n")
output_file.write("};\n\n")
output_file.write("bool HasSideEffects(IROps Op) {\n")
output_file.write(" return SideEffects[Op];\n")
output_file.write("}\n")
output_file.write(f"bool {prop}(IROps Op) {{\n")
output_file.write(f" return {array}[Op];\n")
output_file.write("}\n")
output_file.write("#undef IROP_HASSIDEEFFECTS_IMPL\n")
output_file.write("#endif\n\n")
@@ -627,6 +650,10 @@ def print_ir_allocator_helpers():
output_file.write(") {\n")
# Save NZCV if needed before clobbering NZCV
if op.ImplicitFlagClobber:
output_file.write("\t\tSaveNZCV();")
output_file.write("\t\tauto Op = AllocateOp<IROp_{}, IROps::OP_{}>();\n".format(op.Name, op.Name.upper()))
if op.SSAArgNum != 0:
@@ -675,7 +702,8 @@ def print_ir_allocator_helpers():
output_file.write("\t\t#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
for Validation in op.EmitValidation:
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"\");\n".format(Validation))
Sanitized = Validation.replace("\"", "\\\"")
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("\t\t#endif\n")
output_file.write("\t\treturn Op;\n")
@@ -730,10 +758,38 @@ def print_ir_parser_switch_helper():
output_file.write("#undef IROP_PARSER_SWITCH_HELPERS\n")
output_file.write("#endif\n")
if (len(sys.argv) < 3):
def print_ir_dispatcher_defs():
output_dispatch_file.write("#ifdef IROP_DISPATCH_DEFS\n")
for op in IROps:
if op.Name != "Last" and op.SwitchGen and op.JITDispatch and op.JITDispatchOverride == None:
output_dispatch_file.write("DEF_OP({});\n".format(op.Name))
output_dispatch_file.write("#undef IROP_DISPATCH_DEFS\n")
output_dispatch_file.write("#endif\n")
def print_ir_dispatcher_dispatch():
output_dispatch_file.write("#ifdef IROP_DISPATCH_DISPATCH\n")
for op in IROps:
if op.Name != "Last" and op.JITDispatch:
DispatchName = op.Name
if op.JITDispatchOverride != None:
DispatchName = op.JITDispatchOverride
if (op.DynamicDispatch):
output_dispatch_file.write("REGISTER_OP_RT({}, {});\n".format(op.Name.upper(), DispatchName))
else:
output_dispatch_file.write("REGISTER_OP({}, {});\n".format(op.Name.upper(), DispatchName))
output_dispatch_file.write("#undef IROP_DISPATCH_DISPATCH\n")
output_dispatch_file.write("#endif\n")
if (len(sys.argv) < 4):
ExitError()
output_filename = sys.argv[2]
output_dispatcher_filename = sys.argv[3]
json_file = open(sys.argv[1], "r")
json_text = json_file.read()
json_file.close()
@@ -763,3 +819,10 @@ print_ir_allocator_helpers()
print_ir_parser_switch_helper()
output_file.close()
output_dispatch_file = open(output_dispatcher_filename, "w")
print_ir_dispatcher_defs()
print_ir_dispatcher_dispatch()
output_dispatch_file.close()
+25 -56
View File
@@ -107,9 +107,19 @@ set (SRCS
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp
Interface/Core/JIT/Arm64/Arm64Relocations.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -141,7 +151,7 @@ set (SRCS
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/SyscallOptimization.cpp
Interface/IR/Passes/InlineCallOptimization.cpp
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
@@ -159,24 +169,7 @@ if (ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
Utils/AllocatorOverride.cpp)
endif()
if (ENABLE_INTERPRETER)
list(APPEND SRCS
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp)
endif()
set(DEFINES -DTHREAD_LOCAL=_Thread_local)
set(DEFINES -DTHREAD_LOCAL=_Thread_local -DJIT_ARM64)
if (_M_X86_64)
list(APPEND DEFINES -D_M_X86_64=1)
@@ -195,41 +188,14 @@ if (ENABLE_VIXL_DISASSEMBLER)
list(APPEND DEFINES -DVIXL_DISASSEMBLER=1)
endif()
if (ENABLE_JIT_X86_64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
Interface/Core/JIT/x86_64/ALUOps.cpp
Interface/Core/JIT/x86_64/AtomicOps.cpp
Interface/Core/JIT/x86_64/BranchOps.cpp
Interface/Core/JIT/x86_64/ConversionOps.cpp
Interface/Core/JIT/x86_64/EncryptionOps.cpp
Interface/Core/JIT/x86_64/FlagOps.cpp
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp
Interface/Core/JIT/x86_64/x64Relocations.cpp
)
list(APPEND DEFINES -DJIT_X86_64)
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=__attribute__((preserve_all));-DFEXCORE_HAS_PRESERVE_ALL_ATTR=1")
else()
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=;-DFEXCORE_HAS_PRESERVE_ALL_ATTR=0")
endif()
if (ENABLE_JIT_ARM64)
list(APPEND DEFINES -DJIT_ARM64)
list(APPEND SRCS
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp
Interface/Core/JIT/Arm64/Arm64Relocations.cpp
)
endif()
# Some defines for the softfloat library
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ")
set (LIBS fmt::fmt vixl xxhash FEXHeaderUtils)
@@ -255,15 +221,16 @@ configure_file(
# Generate IR include file
set(OUTPUT_IR_FOLDER "${CMAKE_BINARY_DIR}/include/FEXCore/IR")
set(OUTPUT_NAME "${OUTPUT_IR_FOLDER}/IRDefines.inc")
set(OUTPUT_DISPATCHER_NAME "${OUTPUT_IR_FOLDER}/IRDefines_Dispatch.inc")
set(INPUT_NAME "${CMAKE_CURRENT_SOURCE_DIR}/Interface/IR/IR.json")
file(MAKE_DIRECTORY "${OUTPUT_IR_FOLDER}")
add_custom_command(
OUTPUT "${OUTPUT_NAME}"
OUTPUT "${OUTPUT_NAME}" "${OUTPUT_DISPATCHER_NAME}"
DEPENDS "${INPUT_NAME}"
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py" "${INPUT_NAME}" "${OUTPUT_NAME}"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py" "${INPUT_NAME}" "${OUTPUT_NAME}" "${OUTPUT_DISPATCHER_NAME}"
)
set_source_files_properties(${OUTPUT_NAME} PROPERTIES
@@ -395,6 +362,7 @@ function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
target_link_libraries(${Name} FEXCore_Base)
target_compile_options(${Name} PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
AddDefaultOptionsToTarget(${Name})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
@@ -403,6 +371,7 @@ endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} FEXCore_Base)
target_compile_options(${Name} PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
if (MINGW_BUILD)
# Mingw build isn't building a linux shared library, so it can't have a SONAME.
+1
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/Allocator.h>
+86 -36
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/fextl/fmt.h>
#include "Common/JitSymbols.h"
@@ -26,42 +27,6 @@ namespace FEXCore {
fd = open(PerfMap.c_str(), O_CREAT | O_TRUNC | O_WRONLY | O_APPEND, 0644);
}
void JITSymbols::Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
@@ -86,4 +51,89 @@ namespace FEXCore {
}
}
// Buffered JIT symbols.
void JITSymbols::Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, GuestAddr, CodeSize);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, CodeSize, Name, Offset);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::RegisterNamedRegion(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}\n", HostAddr, CodeSize, Name);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
RegisterNamedRegion(Buffer, HostAddr, CodeSize, Name);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::WriteBuffer(Core::JITSymbolBuffer *Buffer, bool ForceWrite) {
auto Now = std::chrono::steady_clock::now();
if (!ForceWrite) {
if (((Buffer->LastWrite - Now) < Buffer->MAXIMUM_THRESHOLD) &&
Buffer->Offset < Buffer->NEEDS_WRITE_DISTANCE) {
// Still buffering, no need to write.
return;
}
}
Buffer->LastWrite = Now;
auto Result = write(fd, Buffer->Buffer, Buffer->Offset);
if (Result == -1 && errno == EBADF) {
fd = -1;
}
Buffer->Offset = 0;
}
} // namespace FEXCore
+15 -3
View File
@@ -1,5 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <chrono>
#include <cstdint>
#include <cstdio>
#include <memory>
@@ -12,13 +17,20 @@ public:
~JITSymbols();
void InitFile();
void Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
// Allocate JIT buffer.
static fextl::unique_ptr<Core::JITSymbolBuffer> AllocateBuffer() {
return fextl::make_unique<Core::JITSymbolBuffer>();
}
void Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name);
private:
int fd{-1};
void WriteBuffer(Core::JITSymbolBuffer *Buffer, bool ForceWrite = false);
};
}
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
extF80_roundToInt( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
extF80_to_i32( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
extF80_to_i64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t
extF80_to_ui64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t a )
{
union ui128_f128 uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t a )
{
union ui32_f32 uA;
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t a )
{
union ui64_f64 uA;
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t i32_to_extF80( int32_t a )
{
uint_fast16_t uiZ64;
@@ -68,9 +68,11 @@ uint_fast64_t
uint_fast64_t softfloat_roundMToUI64( bool, uint32_t *, uint_fast8_t, bool );
#endif
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t softfloat_roundToI32( bool, uint_fast64_t, uint_fast8_t, bool );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
@@ -109,8 +111,10 @@ float16_t
#define isNaNF32UI( a ) (((~(a) & 0x7F800000) == 0) && ((a) & 0x007FFFFF))
struct exp16_sig32 { int_fast16_t exp; uint_fast32_t sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig32 softfloat_normSubnormalF32Sig( uint_fast32_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t softfloat_roundPackToF32( bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_normRoundPackToF32( bool, int_fast16_t, uint_fast32_t );
@@ -130,8 +134,10 @@ float32_t
#define isNaNF64UI( a ) (((~(a) & UINT64_C( 0x7FF0000000000000 )) == 0) && ((a) & UINT64_C( 0x000FFFFFFFFFFFFF )))
struct exp16_sig64 { int_fast16_t exp; uint_fast64_t sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig64 softfloat_normSubnormalF64Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t softfloat_roundPackToF64( bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_normRoundPackToF64( bool, int_fast16_t, uint_fast64_t );
@@ -155,11 +161,14 @@ float64_t
*----------------------------------------------------------------------------*/
struct exp32_sig64 { int_fast32_t exp; uint64_t sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig64 softfloat_normSubnormalExtF80Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
@@ -181,6 +190,7 @@ extFloat80_t
#define isNaNF128UI( a64, a0 ) (((~(a64) & UINT64_C( 0x7FFF000000000000 )) == 0) && (a0 || ((a64) & UINT64_C( 0x0000FFFFFFFFFFFF ))))
struct exp32_sig128 { int_fast32_t exp; struct uint128 sig; };
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig128
softfloat_normSubnormalF128Sig( uint_fast64_t, uint_fast64_t );
@@ -53,6 +53,7 @@ INLINE
uint64_t softfloat_shortShiftRightJam64( uint64_t a, uint_fast8_t dist )
{ return a>>dist | ((a & (((uint_fast64_t) 1<<dist) - 1)) != 0); }
#else
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shortShiftRightJam64( uint64_t a, uint_fast8_t dist );
#endif
#endif
@@ -74,6 +75,7 @@ INLINE uint32_t softfloat_shiftRightJam32( uint32_t a, uint_fast16_t dist )
(dist < 31) ? a>>dist | ((uint32_t) (a<<(-dist & 31)) != 0) : (a != 0);
}
#else
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_shiftRightJam32( uint32_t a, uint_fast16_t dist );
#endif
#endif
@@ -95,6 +97,7 @@ INLINE uint64_t softfloat_shiftRightJam64( uint64_t a, uint_fast32_t dist )
(dist < 63) ? a>>dist | ((uint64_t) (a<<(-dist & 63)) != 0) : (a != 0);
}
#else
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shiftRightJam64( uint64_t a, uint_fast32_t dist );
#endif
#endif
@@ -148,6 +151,7 @@ INLINE uint_fast8_t softfloat_countLeadingZeros32( uint32_t a )
return count;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros32( uint32_t a );
#endif
#endif
@@ -157,6 +161,7 @@ uint_fast8_t softfloat_countLeadingZeros32( uint32_t a );
| Returns the number of leading 0 bits before the most-significant 1 bit of
| 'a'. If 'a' is zero, 64 is returned.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros64( uint64_t a );
#endif
@@ -178,6 +183,7 @@ extern const uint16_t softfloat_approxRecip_1k1s[16];
#ifdef SOFTFLOAT_FAST_DIV64TO32
#define softfloat_approxRecip32_1( a ) ((uint32_t) (UINT64_C( 0x7FFFFFFFFFFFFFFF ) / (uint32_t) (a)))
#else
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecip32_1( uint32_t a );
#endif
#endif
@@ -204,6 +210,7 @@ extern const uint16_t softfloat_approxRecipSqrt_1k1s[16];
| returned is also always within the range 0.5 to 1; thus, the most-
| significant bit of the result is always set.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecipSqrt32_1( unsigned int oddExpA, uint32_t a );
#endif
@@ -240,6 +247,7 @@ INLINE
bool softfloat_le128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{ return (a64 < b64) || ((a64 == b64) && (a0 <= b0)); }
#else
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_le128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
#endif
@@ -255,6 +263,7 @@ INLINE
bool softfloat_lt128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{ return (a64 < b64) || ((a64 == b64) && (a0 < b0)); }
#else
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_lt128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
#endif
@@ -275,6 +284,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftLeft128( uint64_t a64, uint64_t a0, uint_fast8_t dist );
#endif
@@ -296,6 +306,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftRight128( uint64_t a64, uint64_t a0, uint_fast8_t dist );
#endif
@@ -413,6 +424,7 @@ struct uint64_extra
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint64_extra
softfloat_shiftRightJam64Extra(
uint64_t a, uint64_t extra, uint_fast32_t dist );
@@ -492,6 +504,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_add128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
@@ -528,6 +541,7 @@ struct uint128
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_sub128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 );
#endif
@@ -562,6 +576,7 @@ INLINE struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b )
return z;
}
#else
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b );
#endif
#endif
@@ -570,6 +585,7 @@ struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b );
/*----------------------------------------------------------------------------
| Returns the 128-bit product of 'a' and 'b'.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64To128( uint64_t a, uint64_t b );
#endif
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_add128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_add128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extern const uint16_t softfloat_approxRecip_1k0s[16];
extern const uint16_t softfloat_approxRecip_1k1s[16];
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecip32_1( uint32_t a )
{
int index;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extern const uint16_t softfloat_approxRecipSqrt_1k0s[];
extern const uint16_t softfloat_approxRecipSqrt_1k1s[];
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_approxRecipSqrt32_1( unsigned int oddExpA, uint32_t a )
{
int index;
@@ -44,6 +44,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| floating-point NaN, and returns the bit pattern of this value as an unsigned
| integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr )
{
struct uint128 uiZ;
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| Converts the common NaN pointed to by `aPtr' into a 128-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToF128UI( const struct commonNaN *aPtr )
{
struct uint128 uiZ;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| Converts the common NaN pointed to by `aPtr' into a 32-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast32_t softfloat_commonNaNToF32UI( const struct commonNaN *aPtr )
{
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| Converts the common NaN pointed to by `aPtr' into a 64-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t softfloat_commonNaNToF64UI( const struct commonNaN *aPtr )
{
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#define softfloat_countLeadingZeros32 softfloat_countLeadingZeros32
#include "primitives.h"
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros32( uint32_t a )
{
uint_fast8_t count;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#define softfloat_countLeadingZeros64 softfloat_countLeadingZeros64
#include "primitives.h"
FEXCORE_PRESERVE_ALL_ATTR
uint_fast8_t softfloat_countLeadingZeros64( uint64_t a )
{
uint_fast8_t count;
@@ -46,6 +46,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| location pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
@@ -47,6 +47,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid exception
| is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
@@ -45,6 +45,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| location pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr )
{
@@ -45,6 +45,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| location pointed to by `zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_le128
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_le128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_lt128
FEXCORE_PRESERVE_ALL_ATTR
bool softfloat_lt128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_mul64ByShifted32To128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64ByShifted32To128( uint64_t a, uint32_t b )
{
uint_fast64_t mid;
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_mul64To128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_mul64To128( uint64_t a, uint64_t b )
{
uint32_t a32, a0, b32, b0;
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
bool sign,
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig64 softfloat_normSubnormalExtF80Sig( uint_fast64_t sig )
{
int_fast8_t shiftDist;
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp32_sig128
softfloat_normSubnormalF128Sig( uint_fast64_t sig64, uint_fast64_t sig0 )
{
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig32 softfloat_normSubnormalF32Sig( uint_fast32_t sig )
{
int_fast8_t shiftDist;
@@ -38,6 +38,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "platform.h"
#include "internals.h"
FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig64 softfloat_normSubnormalF64Sig( uint_fast64_t sig )
{
int_fast8_t shiftDist;
@@ -50,6 +50,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| result. If either original floating-point value is a signaling NaN, the
| invalid exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
uint_fast16_t uiA64,
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
bool sign,
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float32_t
softfloat_roundPackToF32( bool sign, int_fast16_t exp, uint_fast32_t sig )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float64_t
softfloat_roundPackToF64( bool sign, int_fast16_t exp, uint_fast64_t sig )
{
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
softfloat_roundToI32(
bool sign, uint_fast64_t sig, uint_fast8_t roundingMode, bool exact )
@@ -41,6 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "specialize.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
bool sign,
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shiftRightJam32
FEXCORE_PRESERVE_ALL_ATTR
uint32_t softfloat_shiftRightJam32( uint32_t a, uint_fast16_t dist )
{
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shiftRightJam64
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shiftRightJam64( uint64_t a, uint_fast32_t dist )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shiftRightJam64Extra
FEXCORE_PRESERVE_ALL_ATTR
struct uint64_extra
softfloat_shiftRightJam64Extra(
uint64_t a, uint64_t extra, uint_fast32_t dist )
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shortShiftLeft128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftLeft128( uint64_t a64, uint64_t a0, uint_fast8_t dist )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shortShiftRight128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_shortShiftRight128( uint64_t a64, uint64_t a0, uint_fast8_t dist )
{
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_shortShiftRightJam64
FEXCORE_PRESERVE_ALL_ATTR
uint64_t softfloat_shortShiftRightJam64( uint64_t a, uint_fast8_t dist )
{
@@ -40,6 +40,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#ifndef softfloat_sub128
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_sub128( uint64_t a64, uint64_t a0, uint64_t b64, uint64_t b0 )
{
@@ -92,6 +92,7 @@ enum {
/*----------------------------------------------------------------------------
| Routine to raise any or all of the software floating-point exception flags.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t );
/*----------------------------------------------------------------------------
@@ -110,6 +111,7 @@ float16_t ui64_to_f16( uint64_t );
float32_t ui64_to_f32( uint64_t );
float64_t ui64_to_f64( uint64_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t ui64_to_extF80( uint64_t );
float128_t ui64_to_f128( uint64_t );
#endif
@@ -119,6 +121,7 @@ float16_t i32_to_f16( int32_t );
float32_t i32_to_f32( int32_t );
float64_t i32_to_f64( int32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t i32_to_extF80( int32_t );
float128_t i32_to_f128( int32_t );
#endif
@@ -183,6 +186,7 @@ int_fast64_t f32_to_i64_r_minMag( float32_t, bool );
float16_t f32_to_f16( float32_t );
float64_t f32_to_f64( float32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t );
float128_t f32_to_f128( float32_t );
#endif
@@ -218,6 +222,7 @@ int_fast64_t f64_to_i64_r_minMag( float64_t, bool );
float16_t f64_to_f16( float64_t );
float32_t f64_to_f32( float64_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t );
float128_t f64_to_f128( float64_t );
#endif
@@ -250,26 +255,41 @@ extern THREAD_LOCAL uint_fast8_t extF80_roundingPrecision;
*----------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_FAST_INT64
uint_fast32_t extF80_to_ui32( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t extF80_to_ui64( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t extF80_to_i32( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t extF80_to_i64( extFloat80_t, uint_fast8_t, bool );
uint_fast32_t extF80_to_ui32_r_minMag( extFloat80_t, bool );
uint_fast64_t extF80_to_ui64_r_minMag( extFloat80_t, bool );
int_fast32_t extF80_to_i32_r_minMag( extFloat80_t, bool );
int_fast64_t extF80_to_i64_r_minMag( extFloat80_t, bool );
float16_t extF80_to_f16( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_roundToInt( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t, extFloat80_t );
bool extF80_le( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t, extFloat80_t );
bool extF80_eq_signaling( extFloat80_t, extFloat80_t );
bool extF80_le_quiet( extFloat80_t, extFloat80_t );
@@ -320,6 +340,7 @@ int_fast64_t f128_to_i64_r_minMag( float128_t, bool );
float16_t f128_to_f16( float128_t );
float32_t f128_to_f32( float128_t );
float64_t f128_to_f64( float128_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t );
float128_t f128_roundToInt( float128_t, uint_fast8_t, bool );
float128_t f128_add( float128_t, float128_t );
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| to substitute a result value. If traps are not implemented, this routine
| should be simply `softfloat_exceptionFlags |= flags;'.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t flags )
{
@@ -135,12 +135,14 @@ uint_fast16_t
| location pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 32-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast32_t softfloat_commonNaNToF32UI( const struct commonNaN *aPtr );
/*----------------------------------------------------------------------------
@@ -170,12 +172,14 @@ uint_fast32_t
| location pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 64-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t softfloat_commonNaNToF64UI( const struct commonNaN *aPtr );
/*----------------------------------------------------------------------------
@@ -215,6 +219,7 @@ uint_fast64_t
| location pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
@@ -224,6 +229,7 @@ void
| floating-point NaN, and returns the bit pattern of this value as an unsigned
| integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr );
/*----------------------------------------------------------------------------
@@ -235,6 +241,7 @@ struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr );
| result. If either original floating-point value is a signaling NaN, the
| invalid exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
uint_fast16_t uiA64,
@@ -264,6 +271,7 @@ struct uint128
| pointed to by 'zPtr'. If the NaN is a signaling NaN, the invalid exception
| is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
@@ -272,6 +280,7 @@ void
| Converts the common NaN pointed to by 'aPtr' into a 128-bit floating-point
| NaN, and returns the bit pattern of this value as an unsigned integer.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
struct uint128 softfloat_commonNaNToF128UI( const struct commonNaN * );
/*----------------------------------------------------------------------------
@@ -39,6 +39,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "internals.h"
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t ui64_to_extF80( uint64_t a )
{
uint_fast16_t uiZ64;
+20
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/BitUtils.h>
@@ -62,6 +63,7 @@ struct FEX_PACKED X80SoftFloat {
}
// Ops
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FADD(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -83,6 +85,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSUB(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -104,6 +107,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FMUL(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -125,6 +129,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FDIV(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -146,6 +151,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FREM(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -168,6 +174,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FREM1(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -190,14 +197,17 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -221,6 +231,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FXTRACT_EXP(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -242,12 +253,14 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static void FCMP(X80SoftFloat const &lhs, X80SoftFloat const &rhs, bool *eq, bool *lt, bool *nan) {
*eq = extF80_eq(lhs, rhs);
*lt = extF80_lt(lhs, rhs);
*nan = IsNan(lhs) || IsNan(rhs);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSCALE(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -276,6 +289,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat F2XM1(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -299,6 +313,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FYL2X(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -324,6 +339,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FATAN(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -349,6 +365,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FTAN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -372,6 +389,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSIN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -394,6 +412,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FCOS(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
@@ -416,6 +435,7 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSQRT(X80SoftFloat const &lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
+1
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/string.h>
+4 -13
View File
@@ -1,5 +1,5 @@
// SPDX-License-Identifier: MIT
#include "Common/StringConv.h"
#include "Common/StringUtils.h"
#include "FEXCore/Utils/EnumUtils.h"
#include <FEXCore/Config/Config.h>
@@ -7,6 +7,7 @@
#include <FEXCore/Utils/CPUInfo.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/StringUtils.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/list.h>
#include <FEXCore/fextl/map.h>
@@ -334,16 +335,11 @@ namespace DefaultValues {
// Sanitize Core option
FEX_CONFIG_OPT(Core, CORE);
#if (_M_X86_64)
constexpr uint32_t MaxCoreNumber = 2;
#else
constexpr uint32_t MaxCoreNumber = 1;
#endif
#ifdef INTERPRETER_ENABLED
constexpr uint32_t MinCoreNumber = 0;
#else
constexpr uint32_t MinCoreNumber = 1;
constexpr uint32_t MaxCoreNumber = 0;
#endif
if (Core > MaxCoreNumber || Core < MinCoreNumber) {
if (Core > MaxCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, fextl::fmt::format("{}", static_cast<uint32_t>(FEXCore::Config::CONFIG_IRJIT)));
}
@@ -352,11 +348,6 @@ namespace DefaultValues {
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION)) {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(Core, CORE);
if (CacheObjectCodeCompilation() && Core() == FEXCore::Config::CONFIG_INTERPRETER) {
// If running the interpreter then disable cache code compilation
FEXCore::Config::Erase(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION);
}
}
fextl::string ContainerPrefix { FindContainerPrefix() };
+28 -4
View File
@@ -6,12 +6,12 @@
"Default": "FEXCore::Config::ConfigCore::CONFIG_IRJIT",
"TextDefault": "irjit",
"ShortArg": "c",
"Choices": [ "irint", "irjit", "host" ],
"Choices": [ "irjit", "host" ],
"ArgumentHandler": "CoreHandler",
"Desc": [
"Which CPU core to use",
"host only exists on x86_64",
"[irint, irjit, host]"
"[irjit, host]"
]
},
"Multiblock": {
@@ -59,6 +59,8 @@
"DISABLESVE": "disablesve",
"ENABLEAVX": "enableavx",
"DISABLEAVX": "disableavx",
"ENABLEAVX2": "enableavx2",
"DISABLEAVX2": "disableavx2",
"ENABLEAFP": "enableafp",
"DISABLEAFP": "disableafp",
"ENABLELRCPC": "enablelrcpc",
@@ -76,13 +78,22 @@
"ENABLEATOMICS": "enableatomics",
"DISABLEATOMICS": "disableatomics",
"ENABLEFCMA": "enablefcma",
"DISABLEFCMA": "disablefcma"
"DISABLEFCMA": "disablefcma",
"ENABLEFLAGM": "enableflagm",
"DISABLEFLAGM": "disableflagm",
"ENABLEFLAGM2": "enableflagm2",
"DISABLEFLAGM2": "disableflagm2",
"ENABLECRYPTO": "enablecrypto",
"DISABLECRYPTO": "disablecrypto",
"ENABLERPRES": "enablerpres",
"DISABLERPRES": "disablerpres"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
"\toff: Default CPU features queried from CPU features",
"\t{enable,disable}sve: Will force enable or disable sve even if the host doesn't support it",
"\t{enable,disable}avx: Will force enable or disable avx even if the host doesn't support it",
"\t{enable,disable}avx2: Will force enable or disable avx2 even if the host doesn't support it",
"\t{enable,disable}afp: Will force enable or disable afp even if the host doesn't support it",
"\t{enable,disable}lrcpc: Will force enable or disable lrcpc even if the host doesn't support it",
"\t{enable,disable}lrcpc2: Will force enable or disable lrcpc2 even if the host doesn't support it",
@@ -91,7 +102,11 @@
"\t{enable,disable}rng: Will force enable or disable rng even if the host doesn't support it",
"\t{enable,disable}clzero: Will force enable or disable clzero even if the host doesn't support it",
"\t{enable,disable}atomics: Will force enable or disable ARMv8.1 LSE atomics even if the host doesn't support it",
"\t{enable,disable}fcma: Will force enable or disable fcma even if the host doesn't support it"
"\t{enable,disable}fcma: Will force enable or disable fcma even if the host doesn't support it",
"\t{enable,disable}flagm: Will force enable or disable flagm even if the host doesn't support it",
"\t{enable,disable}flagm2: Will force enable or disable flagm2 even if the host doesn't support it",
"\t{enable,disable}crypto: Will force enable or disable crypto extensions even if the host doesn't support it",
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it"
]
}
},
@@ -490,6 +505,15 @@
"IS64BIT_MODE": {
"Type": "bool",
"Default": "false"
},
"DISABLE_VIXL_INDIRECT_RUNTIME_CALLS": {
"Type": "bool",
"Default": "true",
"Desc": [
"This option is used for the InstructionCountCI so it can generate the same codegen between Arm64 hosts and vixl simulator hosts.",
"Vixl simulator indirect runtime calls are a special hlt instruction with metadata after it. Effectively making a custom call instruction.",
"With visual simulator calls disabled, the code generation would be the same as on a native Arm64 host, but running the code is broken."
]
}
}
}
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
@@ -47,6 +48,10 @@ namespace FEXCore::Context {
CompileBlock(Thread->CurrentFrame, GuestRIP);
}
void FEXCore::Context::ContextImpl::CompileRIPCount(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst) {
CompileBlock(Thread->CurrentFrame, GuestRIP, MaxInst);
}
FEXCore::Context::ExitReason FEXCore::Context::ContextImpl::GetExitReason() {
return ParentThread->ExitReason;
}
+34 -9
View File
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Common/JitSymbols.h"
@@ -45,7 +46,6 @@ namespace CodeSerialize {
namespace CPU {
class Arm64JITCore;
class X86JITCore;
class InterpreterCore;
class Dispatcher;
}
namespace HLE {
@@ -87,7 +87,10 @@ namespace FEXCore::Context {
ExitReason RunUntilExit() override;
void ExecuteThread(FEXCore::Core::InternalThreadState *Thread) override;
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
int GetProgramStatus() const override;
@@ -204,7 +207,6 @@ namespace FEXCore::Context {
friend class FEXCore::CPU::X86JITCore;
#endif
friend class FEXCore::CPU::InterpreterCore;
friend class FEXCore::IR::Validation::IRValidation;
struct {
@@ -214,6 +216,9 @@ namespace FEXCore::Context {
// this is for internal use
bool ValidateIRarser { false };
// Used if the JIT needs to have its interrupt fault code emitted.
bool NeedsPendingInterruptFaultCheck { false };
FEX_CONFIG_OPT(Multiblock, MULTIBLOCK);
FEX_CONFIG_OPT(SingleStepConfig, SINGLESTEP);
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
@@ -240,6 +245,7 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
} Config;
FEXCore::HostFeatures HostFeatures;
@@ -279,9 +285,9 @@ namespace FEXCore::Context {
~ContextImpl();
bool IsPaused() const { return !Running; }
void WaitForThreadsToRun();
void WaitForThreadsToRun() override;
void Stop(bool IgnoreCurrentThread);
void WaitForIdle();
void WaitForIdle() override;
void SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event);
bool GetGdbServerStatus() const { return DebugServer != nullptr; }
@@ -320,7 +326,7 @@ namespace FEXCore::Context {
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]] GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo);
[[nodiscard]] GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst);
struct CompileCodeResult {
void* CompiledCode;
@@ -331,8 +337,8 @@ namespace FEXCore::Context {
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]] CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
[[nodiscard]] CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP, uint64_t MaxInst = 0);
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
@@ -347,8 +353,6 @@ namespace FEXCore::Context {
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
fextl::vector<FEXCore::Core::InternalThreadState*>* GetThreads() { return &Threads; }
uint8_t GetGPRSize() const { return Config.Is64BitMode ? 8 : 4; }
FEXCore::JITSymbols Symbols;
@@ -378,10 +382,31 @@ namespace FEXCore::Context {
UpdateAtomicTSOEmulationConfig();
}
// Returns if Software TSO emulation is required.
// NOTE: This doesn't necessary return if Atomic-based TSO is currently enabled.
// This will still return true if on a single thread and TSO is currently disabled.
//
// This is to ensure that if early initialization checks CPU features and TSO /could/ be enabled, that
// we return consistent results.
//
// To check if Atomic TSO is currently enabled in the JIT, use `IsAtomicTSOEnabled` instead.
bool SoftwareTSORequired() const {
if (SupportsHardwareTSO) return false;
return Config.TSOEnabled;
}
void EnableExitOnHLT() override { ExitOnHLT = true; }
bool ExitOnHLTEnabled() const { return ExitOnHLT; }
ThreadsState GetThreads() override {
return ThreadsState {
.ParentThread = ParentThread,
.Threads = &Threads,
};
}
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
protected:
@@ -1,6 +1,8 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include "Interface/HLE/Thunks/Thunks.h"
@@ -66,11 +68,102 @@ namespace x64 {
};
// v8..v15 = (lower 64bits) Callee saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 12> RAFPR = {
// v0 ~ v3 are used as temps.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> RAFPR = {
// v0 ~ v1 are used as temps.
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
// FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
};
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 7> PreserveAll_SRA = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8,
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
};
constexpr uint32_t PreserveAll_SRAMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRA) {
switch (Reg.Idx()) {
case 0:
case 1:
case 2:
case 3:
case 4:
case 5:
case 6:
case 7:
case 8:
case 16:
case 17:
Mask |= (1U << Reg.Idx());
break;
default: break;
}
}
return Mask;
}()
};
// Dynamic GPRs
constexpr std::array<FEXCore::ARMEmitter::Register, 1> PreserveAll_Dynamic = {
// Only LR needs to get saved.
FEXCore::ARMEmitter::Reg::r30
};
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
// None.
};
constexpr uint32_t PreserveAll_SRAFPRMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPR) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs
// - v0-v7
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
// v0 ~ v1 are temps
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
};
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
// This is /all/ of the SRA registers
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPRSVE) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs when the host supports SVE-256bit.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> PreserveAll_DynamicFPRSVE = {
// v0 ~ v1 are used as temps.
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
@@ -127,11 +220,106 @@ namespace x32 {
};
// v8..v15 = (lower 64bits) Callee saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 20> RAFPR = {
// v0 ~ v3 are used as temps.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> RAFPR = {
// v0 ~ v1 are used as temps.
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
// FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25,
FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29,
FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31
};
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 5> PreserveAll_SRA = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8,
};
constexpr uint32_t PreserveAll_SRAMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRA) {
switch (Reg.Idx()) {
case 0:
case 1:
case 2:
case 3:
case 4:
case 5:
case 6:
case 7:
case 8:
case 16:
case 17:
Mask |= (1U << Reg.Idx());
break;
default: break;
}
}
return Mask;
}()
};
// Dynamic GPRs
constexpr std::array<FEXCore::ARMEmitter::Register, 3> PreserveAll_Dynamic = {
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
FEXCore::ARMEmitter::Reg::r30
};
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
// None.
};
constexpr uint32_t PreserveAll_SRAFPRMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPR) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs
// - v0-v7
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
// v0 ~ v1 are temps
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
};
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
// This is /all/ of the SRA registers
constexpr std::array<FEXCore::ARMEmitter::VRegister, 8> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {
[]() -> uint32_t {
uint32_t Mask{};
for (auto Reg : PreserveAll_SRAFPRSVE) {
Mask |= (1U << Reg.Idx());
}
return Mask;
}()
};
// Dynamic FPRs when the host supports SVE-256bit.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> PreserveAll_DynamicFPRSVE = {
// v0 ~ v1 are used as temps.
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
@@ -147,8 +335,8 @@ namespace x32 {
}
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, size_t size)
: Emitter(size ? (uint8_t*)FEXCore::Allocator::VirtualAlloc(size, true) : nullptr, size)
Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, void* EmissionPtr, size_t size)
: Emitter(static_cast<uint8_t*>(EmissionPtr), size)
, EmitterCTX {ctx}
#ifdef VIXL_SIMULATOR
, Simulator {&SimDecoder}
@@ -191,13 +379,6 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, size_t size)
}
}
Arm64Emitter::~Arm64Emitter() {
auto BufferSize = GetBufferSize();
if (BufferSize) {
FEXCore::Allocator::VirtualFree(GetBufferBase(), BufferSize);
}
}
void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad) {
bool Is64Bit = s == ARMEmitter::Size::i64Bit;
int Segments = Is64Bit ? 4 : 2;
@@ -393,6 +574,24 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Disable AFP features when spilling registers.
//
// Disable FPCR.NEP and FPCR.AH
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
//
// Additional interesting AFP bits:
// FIZ(0): Flush Inputs to Zero
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
bic(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
}
#endif
if (!StaticRegisterAllocation()) {
return;
}
@@ -457,6 +656,38 @@ void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FP
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask) {
FEXCore::ARMEmitter::Register TmpReg = FEXCore::ARMEmitter::Reg::r0;
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 1 GPR for a temp");
bool FoundRegister{};
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & GPRFillMask)) {
TmpReg = Reg;
FoundRegister = true;
break;
}
}
LOGMAN_THROW_A_FMT(FoundRegister, "Didn't have an SRA register to use as a temporary while spilling!");
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Enable AFP features when filling JIT state.
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 1 GPR for a temp");
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
// Enable FPCR.NEP and FPCR.AH
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
//
// Additional interesting AFP bits:
// FIZ(0): Flush Inputs to Zero
orr(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
}
#endif
if (!StaticRegisterAllocation()) {
return;
}
@@ -467,11 +698,11 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
if (EmitterCTX->HostFeatures.SupportsSVE) {
ptrue<ARMEmitter::SubRegSize::i8Bit>(PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
if (EmitterCTX->HostFeatures.SupportsAVX) {
ptrue<ARMEmitter::SubRegSize::i8Bit>(PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
@@ -484,8 +715,6 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
if (GPRFillMask && FPRFillMask == ~0U) {
// Optimize the common case where we can fill four registers per instruction.
// Use one of the filling static registers before we fill it.
auto TmpReg = StaticRegisters[FindFirstSetBit(GPRFillMask)];
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
@@ -532,6 +761,107 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
}
}
void Arm64Emitter::PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
if (SVERegs) {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
st2b(Reg1.Z(), Reg2.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 2);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
st4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 4);
}
}
else {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), TmpReg, 32);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), TmpReg, 64);
}
}
}
void Arm64Emitter::PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs) {
size_t i = 0;
for (; i < (Regs.size() % 2); ++i) {
const auto Reg1 = Regs[i];
str<ARMEmitter::IndexType::POST>(Reg1.X(), TmpReg, 16);
}
for (; i < Regs.size(); i += 2) {
const auto Reg1 = Regs[i];
const auto Reg2 = Regs[i + 1];
stp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), TmpReg, 16);
}
}
void Arm64Emitter::PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
if (SVERegs) {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
ld2b(Reg1.Z(), Reg2.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 2);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
ld4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 4);
}
} else {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), ARMEmitter::Reg::rsp, 32);
}
for (; i < VRegs.size(); i += 4) {
const auto Reg1 = VRegs[i];
const auto Reg2 = VRegs[i + 1];
const auto Reg3 = VRegs[i + 2];
const auto Reg4 = VRegs[i + 3];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), ARMEmitter::Reg::rsp, 64);
}
}
}
void Arm64Emitter::PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs) {
size_t i = 0;
for (; i < (Regs.size() % 2); ++i) {
const auto Reg1 = Regs[i];
ldr<ARMEmitter::IndexType::POST>(Reg1.X(), ARMEmitter::Reg::rsp, 16);
}
for (; i < Regs.size(); i += 2) {
const auto Reg1 = Regs[i];
const auto Reg2 = Regs[i + 1];
ldp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), ARMEmitter::Reg::rsp, 16);
}
}
void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto GPRSize = (ConfiguredDynamicRegisterBase.size() + 1) * Core::CPUState::GPR_REG_SIZE;
@@ -545,31 +875,13 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
// rsp capable move
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
if (CanUseSVE) {
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
st4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 4);
}
} else {
LOGMAN_THROW_A_FMT(GeneralFPRegisters.size() % 4 == 0, "Needs to have multiple of 4 FPRs for RA");
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), TmpReg, 64);
}
}
LOGMAN_THROW_A_FMT(GeneralFPRegisters.size() % 2 == 0, "Needs to have multiple of 2 FPRs for RA");
for (size_t i = 0; i < ConfiguredDynamicRegisterBase.size(); i += 2) {
const auto Reg1 = ConfiguredDynamicRegisterBase[i];
const auto Reg2 = ConfiguredDynamicRegisterBase[i + 1];
stp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), TmpReg, 16);
}
// Push the vector registers
PushVectorRegisters(TmpReg, CanUseSVE, GeneralFPRegisters);
// Push the general registers.
PushGeneralRegisters(TmpReg, ConfiguredDynamicRegisterBase);
str(ARMEmitter::XReg::lr, TmpReg, 0);
}
@@ -577,34 +889,107 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
void Arm64Emitter::PopDynamicRegsAndLR() {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
if (CanUseSVE) {
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
ld4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 4);
}
} else {
for (size_t i = 0; i < GeneralFPRegisters.size(); i += 4) {
const auto Reg1 = GeneralFPRegisters[i];
const auto Reg2 = GeneralFPRegisters[i + 1];
const auto Reg3 = GeneralFPRegisters[i + 2];
const auto Reg4 = GeneralFPRegisters[i + 3];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), ARMEmitter::Reg::rsp, 64);
}
}
// Pop vectors first
PopVectorRegisters(CanUseSVE, GeneralFPRegisters);
for (size_t i = 0; i < ConfiguredDynamicRegisterBase.size(); i += 2) {
const auto Reg1 = ConfiguredDynamicRegisterBase[i];
const auto Reg2 = ConfiguredDynamicRegisterBase[i + 1];
ldp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), ARMEmitter::Reg::rsp, 16);
}
// Pop GPRs second
PopGeneralRegisters(ConfiguredDynamicRegisterBase);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
}
void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs{};
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs{};
uint32_t PreserveSRAMask{};
uint32_t PreserveSRAFPRMask{};
if (EmitterCTX->Config.Is64BitMode()) {
DynamicGPRs = x64::PreserveAll_Dynamic;
DynamicFPRs = x64::PreserveAll_DynamicFPR;
PreserveSRAMask = x64::PreserveAll_SRAMask;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
}
}
else {
DynamicGPRs = x32::PreserveAll_Dynamic;
DynamicFPRs = x32::PreserveAll_DynamicFPR;
PreserveSRAMask = x32::PreserveAll_SRAMask;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
}
}
const auto GPRSize = AlignUp(DynamicGPRs.size(), 2) * Core::CPUState::GPR_REG_SIZE;
const auto FPRSize = DynamicFPRs.size() * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
// Spill the static registers.
SpillStaticRegs(TmpReg, true, PreserveSRAMask, PreserveSRAFPRMask);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
// rsp capable move
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
// Push the vector registers.
PushVectorRegisters(TmpReg, CanUseSVE, DynamicFPRs);
// Push the general registers.
PushGeneralRegisters(TmpReg, DynamicGPRs);
}
void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs{};
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs{};
uint32_t PreserveSRAMask{};
uint32_t PreserveSRAFPRMask{};
if (EmitterCTX->Config.Is64BitMode()) {
DynamicGPRs = x64::PreserveAll_Dynamic;
DynamicFPRs = x64::PreserveAll_DynamicFPR;
PreserveSRAMask = x64::PreserveAll_SRAMask;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
}
}
else {
DynamicGPRs = x32::PreserveAll_Dynamic;
DynamicFPRs = x32::PreserveAll_DynamicFPR;
PreserveSRAMask = x32::PreserveAll_SRAMask;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
}
}
// Fill the static registers.
FillStaticRegs(true, PreserveSRAMask, PreserveSRAFPRMask);
// Pop the vector registers.
PopVectorRegisters(CanUseSVE, DynamicFPRs);
// Pop the general registers.
PopGeneralRegisters(DynamicGPRs);
}
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetCursorAddress<uint64_t>();
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
@@ -1,10 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include <aarch64/assembler-aarch64.h>
@@ -28,6 +28,10 @@
#include <utility>
#include <span>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
constexpr auto STATE = FEXCore::ARMEmitter::XReg::x28;
@@ -42,8 +46,6 @@ constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x3;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v0;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
constexpr auto VTMP3 = FEXCore::ARMEmitter::VReg::v2;
constexpr auto VTMP4 = FEXCore::ARMEmitter::VReg::v3;
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
@@ -55,8 +57,7 @@ constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_32B = FEXCore::ARMEmitter::PRe
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public FEXCore::ARMEmitter::Emitter {
protected:
Arm64Emitter(FEXCore::Context::ContextImpl *ctx, size_t size);
~Arm64Emitter();
Arm64Emitter(FEXCore::Context::ContextImpl *ctx, void* EmissionPtr = nullptr, size_t size = 0);
FEXCore::Context::ContextImpl *EmitterCTX;
vixl::aarch64::CPU CPU;
@@ -99,12 +100,52 @@ protected:
// We can't guarantee only the lower 64bits are used so flush everything
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
// Generic push and pop vector registers.
void PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs);
void PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs);
void PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg);
void PopDynamicRegsAndLR();
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
// Spills and fills SRA/Dynamic registers that are required for Arm64 `preserve_all` ABI.
// This ABI changes most registers to be callee saved.
// Caller Saved:
// - X0-X8, X16-X18.
// - v0-v7
// - For 256-bit SVE hosts: top 128-bits of v8-v31
//
// Callee Saved:
// - X9-X15, X19-X31
// - Low 128-bits of v8-v31
void SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true);
void FillForPreserveAllABICall(bool FPRs = true);
void SpillForABICall(bool SupportsPreserveAllABI, FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true) {
if (SupportsPreserveAllABI) {
SpillForPreserveAllABICall(TMP1, true);
}
else {
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
}
}
void FillForABICall(bool SupportsPreserveAllABI, bool FPRs = true) {
if (SupportsPreserveAllABI) {
FillForPreserveAllABICall(true);
}
else {
PopDynamicRegsAndLR();
FillStaticRegs();
}
}
void Align16B();
#ifdef VIXL_SIMULATOR
@@ -171,7 +212,15 @@ protected:
// Call type
dc32(vixl::aarch64::kCallRuntime);
}
#else
template<typename R, typename... P>
void GenerateRuntimeCall(R (*Function)(P...)) {
// Explicitly doing nothing.
}
template<typename R, typename... P>
void GenerateIndirectRuntimeCall(ARMEmitter::Register Reg) {
// Explicitly doing nothing.
}
#endif
#ifdef VIXL_SIMULATOR
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* ALU instruction emitters.
*
* Almost all of these operations have `ARMEmitter::Size` as their first argument.
@@ -765,6 +766,11 @@ public:
EvaluateIntoFlags(Op, 1, rn);
}
void cfinv() {
constexpr uint32_t Op = 0b1101'0101'0000'0000'0100'0000'0001'1111;
dc32(Op);
}
// Conditional compare - register
void ccmn(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::StatusFlags flags, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0011'1010'010 << 21;
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* ASIMD instruction emitters.
*
* This contains emitters for vector operations explicitly.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* Branch instruction emitters.
*
* Most of these instructions will use `BackwardLabel`, `ForwardLabel`, or `BiDirectionLabel` to determine where a branch targets.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef>
#include <cstdint>
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* Load-store instruction emitters
*
* For GPR load-stores that take a `Size` argument as their first argument can be 32-bit or 64-bit.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/EnumUtils.h>
@@ -20,12 +21,11 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const Register&, const Register&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
WRegister W() const;
XRegister X() const;
constexpr WRegister W() const;
constexpr XRegister X() const;
private:
uint32_t Index;
@@ -45,16 +45,15 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const WRegister&, const WRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator Register() const {
constexpr operator Register() const {
return Register(Index);
}
XRegister X() const;
Register R() const;
constexpr XRegister X() const;
constexpr Register R() const;
private:
uint32_t Index;
@@ -74,16 +73,15 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const XRegister&, const XRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator Register() const {
constexpr operator Register() const {
return Register(Index);
}
WRegister W() const;
Register R() const;
constexpr WRegister W() const;
constexpr Register R() const;
private:
uint32_t Index;
@@ -92,27 +90,27 @@ namespace FEXCore::ARMEmitter {
static_assert(std::is_trivial_v<Register>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<Register>, "Needs to be standard");
inline WRegister Register::W() const {
inline constexpr WRegister Register::W() const {
return WRegister{Index};
}
inline XRegister Register::X() const {
inline constexpr XRegister Register::X() const {
return XRegister{Index};
}
inline XRegister WRegister::X() const {
inline constexpr XRegister WRegister::X() const {
return XRegister{Index};
}
inline Register WRegister::R() const {
inline constexpr Register WRegister::R() const {
return *this;
}
inline WRegister XRegister::W() const {
inline constexpr WRegister XRegister::W() const {
return WRegister{Index};
}
inline Register XRegister::R() const {
inline constexpr Register XRegister::R() const {
return *this;
}
@@ -259,7 +257,6 @@ namespace FEXCore::ARMEmitter {
class QRegister;
class ZRegister;
/* Unsized ASIMD register class
* This class doesn't imply a size when used, nor implies Vector or Scalar.
* It does imply that this instruction isn't using the register for SVE.
@@ -272,16 +269,16 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const VRegister&, const VRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
BRegister B() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -301,20 +298,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const BRegister&, const BRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister () const {
constexpr operator VRegister() const {
return VRegister(Index);
}
BRegister V() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr BRegister V() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -334,20 +330,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const HRegister&, const HRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister() const {
constexpr operator VRegister() const {
return VRegister(Index);
}
HRegister V() const;
BRegister B() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr HRegister V() const;
constexpr BRegister B() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -367,20 +362,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const SRegister&, const SRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister() const {
constexpr operator VRegister() const {
return VRegister(Index);
}
SRegister V() const;
BRegister B() const;
HRegister H() const;
DRegister D() const;
QRegister Q() const;
ZRegister Z() const;
constexpr SRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -401,20 +395,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const DRegister&, const DRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister() const {
constexpr operator VRegister() const {
return VRegister(Index);
}
DRegister V() const;
BRegister B() const;
HRegister H() const;
SRegister S() const;
QRegister Q() const;
ZRegister Z() const;
constexpr DRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr QRegister Q() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -435,20 +428,19 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const QRegister&, const QRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator VRegister () const {
constexpr operator VRegister() const {
return VRegister(Index);
}
QRegister V() const;
BRegister B() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
ZRegister Z() const;
constexpr QRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr ZRegister Z() const;
private:
uint32_t Index;
@@ -468,16 +460,16 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const ZRegister&, const ZRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
VRegister V() const;
BRegister B() const;
HRegister H() const;
SRegister S() const;
DRegister D() const;
QRegister Q() const;
constexpr VRegister V() const;
constexpr BRegister B() const;
constexpr HRegister H() const;
constexpr SRegister S() const;
constexpr DRegister D() const;
constexpr QRegister Q() const;
private:
uint32_t Index;
@@ -487,142 +479,142 @@ namespace FEXCore::ARMEmitter {
static_assert(std::is_standard_layout_v<ZRegister>, "Needs to be standard");
// VRegister
inline BRegister VRegister::B() const {
inline constexpr BRegister VRegister::B() const {
return BRegister{Index};
}
inline HRegister VRegister::H() const {
inline constexpr HRegister VRegister::H() const {
return HRegister{Index};
}
inline SRegister VRegister::S() const {
inline constexpr SRegister VRegister::S() const {
return SRegister{Index};
}
inline DRegister VRegister::D() const {
inline constexpr DRegister VRegister::D() const {
return DRegister{Index};
}
inline QRegister VRegister::Q() const {
inline constexpr QRegister VRegister::Q() const {
return QRegister{Index};
}
inline ZRegister VRegister::Z() const {
inline constexpr ZRegister VRegister::Z() const {
return ZRegister{Index};
}
// BRegister
inline BRegister BRegister::V() const {
inline constexpr BRegister BRegister::V() const {
return *this;
}
inline HRegister BRegister::H() const {
inline constexpr HRegister BRegister::H() const {
return HRegister{Index};
}
inline SRegister BRegister::S() const {
inline constexpr SRegister BRegister::S() const {
return SRegister{Index};
}
inline DRegister BRegister::D() const {
inline constexpr DRegister BRegister::D() const {
return DRegister{Index};
}
inline QRegister BRegister::Q() const {
inline constexpr QRegister BRegister::Q() const {
return QRegister{Index};
}
inline ZRegister BRegister::Z() const {
inline constexpr ZRegister BRegister::Z() const {
return ZRegister{Index};
}
// HRegister
inline HRegister HRegister::V() const {
inline constexpr HRegister HRegister::V() const {
return *this;
}
inline BRegister HRegister::B() const {
inline constexpr BRegister HRegister::B() const {
return BRegister{Index};
}
inline SRegister HRegister::S() const {
inline constexpr SRegister HRegister::S() const {
return SRegister{Index};
}
inline DRegister HRegister::D() const {
inline constexpr DRegister HRegister::D() const {
return DRegister{Index};
}
inline QRegister HRegister::Q() const {
inline constexpr QRegister HRegister::Q() const {
return QRegister{Index};
}
inline ZRegister HRegister::Z() const {
inline constexpr ZRegister HRegister::Z() const {
return ZRegister{Index};
}
// SRegister
inline SRegister SRegister::V() const {
inline constexpr SRegister SRegister::V() const {
return *this;
}
inline BRegister SRegister::B() const {
inline constexpr BRegister SRegister::B() const {
return BRegister{Index};
}
inline HRegister SRegister::H() const {
inline constexpr HRegister SRegister::H() const {
return HRegister{Index};
}
inline DRegister SRegister::D() const {
inline constexpr DRegister SRegister::D() const {
return DRegister{Index};
}
inline QRegister SRegister::Q() const {
inline constexpr QRegister SRegister::Q() const {
return QRegister{Index};
}
inline ZRegister SRegister::Z() const {
inline constexpr ZRegister SRegister::Z() const {
return ZRegister{Index};
}
// DRegister
inline DRegister DRegister::V() const {
inline constexpr DRegister DRegister::V() const {
return DRegister{Index};
}
inline BRegister DRegister::B() const {
inline constexpr BRegister DRegister::B() const {
return BRegister{Index};
}
inline HRegister DRegister::H() const {
inline constexpr HRegister DRegister::H() const {
return HRegister{Index};
}
inline SRegister DRegister::S() const {
inline constexpr SRegister DRegister::S() const {
return SRegister{Index};
}
inline QRegister DRegister::Q() const {
inline constexpr QRegister DRegister::Q() const {
return QRegister{Index};
}
inline ZRegister DRegister::Z() const {
inline constexpr ZRegister DRegister::Z() const {
return ZRegister{Index};
}
// QRegister
inline QRegister QRegister::V() const {
inline constexpr QRegister QRegister::V() const {
return *this;
}
inline BRegister QRegister::B() const {
inline constexpr BRegister QRegister::B() const {
return BRegister{Index};
}
inline HRegister QRegister::H() const {
inline constexpr HRegister QRegister::H() const {
return HRegister{Index};
}
inline SRegister QRegister::S() const {
inline constexpr SRegister QRegister::S() const {
return SRegister{Index};
}
inline DRegister QRegister::D() const {
inline constexpr DRegister QRegister::D() const {
return DRegister{Index};
}
inline ZRegister QRegister::Z() const {
inline constexpr ZRegister QRegister::Z() const {
return ZRegister{Index};
}
// ZRegister
inline VRegister ZRegister::V() const {
inline constexpr VRegister ZRegister::V() const {
return VRegister(Index);
}
inline BRegister ZRegister::B() const {
inline constexpr BRegister ZRegister::B() const {
return BRegister(Index);
}
inline HRegister ZRegister::H() const {
inline constexpr HRegister ZRegister::H() const {
return HRegister(Index);
}
inline SRegister ZRegister::S() const {
inline constexpr SRegister ZRegister::S() const {
return SRegister(Index);
}
inline DRegister ZRegister::D() const {
inline constexpr DRegister ZRegister::D() const {
return DRegister(Index);
}
inline QRegister ZRegister::Q() const {
inline constexpr QRegister ZRegister::Q() const {
return QRegister(Index);
}
@@ -879,36 +871,28 @@ namespace FEXCore::ARMEmitter {
}
// Zero-cost FPR->GPR
inline
Register ToReg(HRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(HRegister Reg) {
return Register(Reg.Idx());
}
inline
Register ToReg(SRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(SRegister Reg) {
return Register(Reg.Idx());
}
inline
Register ToReg(DRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(DRegister Reg) {
return Register(Reg.Idx());
}
inline
Register ToReg(VRegister Reg) {
return static_cast<Register>(Reg.Idx());
inline constexpr Register ToReg(VRegister Reg) {
return Register(Reg.Idx());
}
// Zero-cost GPR->FPR
inline
VRegister ToVReg(Register Reg) {
return static_cast<VRegister>(Reg.Idx());
inline constexpr VRegister ToVReg(Register Reg) {
return VRegister(Reg.Idx());
}
inline
VRegister ToVReg(XRegister Reg) {
return static_cast<VRegister>(Reg.Idx());
inline constexpr VRegister ToVReg(XRegister Reg) {
return VRegister(Reg.Idx());
}
inline
VRegister ToVReg(WRegister Reg) {
return static_cast<VRegister>(Reg.Idx());
inline constexpr VRegister ToVReg(WRegister Reg) {
return VRegister(Reg.Idx());
}
class PRegisterZero;
@@ -925,12 +909,12 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const PRegister&, const PRegister&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
PRegisterZero Zeroing() const;
PRegisterMerge Merging() const;
constexpr PRegisterZero Zeroing() const;
constexpr PRegisterMerge Merging() const;
private:
uint32_t Index;
@@ -948,14 +932,17 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const PRegisterZero&, const PRegisterZero&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator PRegister() const;
PRegister P() const;
PRegisterMerge Merging() const;
constexpr operator PRegister() const {
return PRegister(Index);
}
constexpr PRegister P() const {
return PRegister(Index);
}
constexpr PRegisterMerge Merging() const;
private:
uint32_t Index;
@@ -973,14 +960,17 @@ namespace FEXCore::ARMEmitter {
friend constexpr auto operator<=>(const PRegisterMerge&, const PRegisterMerge&) = default;
uint32_t Idx() const {
constexpr uint32_t Idx() const {
return Index;
}
operator PRegister() const;
PRegister P() const;
PRegisterZero Zeroing() const;
constexpr operator PRegister() const {
return PRegister(Index);
}
constexpr PRegister P() const {
return PRegister(Index);
}
constexpr PRegisterZero Zeroing() const;
private:
uint32_t Index;
@@ -989,39 +979,21 @@ namespace FEXCore::ARMEmitter {
static_assert(std::is_trivial_v<PRegisterZero>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<PRegisterZero>, "Needs to be standard");
// PRegister
inline PRegisterZero PRegister::Zeroing() const {
inline constexpr PRegisterZero PRegister::Zeroing() const {
return PRegisterZero(Idx());
}
inline PRegisterMerge PRegister::Merging() const {
inline constexpr PRegisterMerge PRegister::Merging() const {
return PRegisterMerge(Idx());
}
// PRegisterZero
inline PRegisterZero::operator PRegister() const {
return PRegister(Index);
}
inline PRegister PRegisterZero::P() const {
return PRegister(Idx());
}
inline PRegisterMerge PRegisterZero::Merging() const {
inline constexpr PRegisterMerge PRegisterZero::Merging() const {
return PRegisterMerge(Idx());
}
// PRegisterMerge
inline PRegisterMerge::operator PRegister() const {
return PRegisterZero(Index);
}
inline PRegister PRegisterMerge::P() const {
return PRegister(Idx());
}
inline PRegisterZero PRegisterMerge::Zeroing() const {
inline constexpr PRegisterZero PRegisterMerge::Zeroing() const {
return PRegisterZero(Idx());
}
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* SVE instruction emitters
* These contain instruction emitters for AArch64 SVE and SVE2 operations.
*
@@ -1383,12 +1384,10 @@ public:
}
// SVE predicate initialize
template <SubRegSize size>
void ptrue(PRegister pd, PredicatePattern pattern) {
void ptrue(SubRegSize size, PRegister pd, PredicatePattern pattern) {
SVEPredicateMisc(0b1000, 0b10000, FEXCore::ToUnderlying(pattern), size, pd);
}
template <SubRegSize size>
void ptrues(PRegister pd, PredicatePattern pattern) {
void ptrues(SubRegSize size, PRegister pd, PredicatePattern pattern) {
SVEPredicateMisc(0b1001, 0b10000, FEXCore::ToUnderlying(pattern), size, pd);
}
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* Scalar instruction emitters.
*
* These contain instruction emitters for scalar ASIMD operations explicitly.
@@ -797,6 +798,52 @@ public:
// XXX:
//
// Floating-point data-processing (1 source)
void fmov(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b000000, rd, rn);
}
void fabs(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b000001, rd, rn);
}
void fneg(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b000010, rd, rn);
}
void fsqrt(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b000011, rd, rn);
}
void frintn(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b001000, rd, rn);
}
void frintp(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b001001, rd, rn);
}
void frintm(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b001010, rd, rn);
}
void frintz(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b001011, rd, rn);
}
void frinta(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b001100, rd, rn);
}
void frintx(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b001110, rd, rn);
}
void frinti(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b001111, rd, rn);
}
void frint32z(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b010000, rd, rn);
}
void frint32x(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b010001, rd, rn);
}
void frint64z(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b010010, rd, rn);
}
void frint64x(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b010011, rd, rn);
}
void fmov(SRegister rd, SRegister rn) {
Float1Source(0, 0, 0b00, 0b000000, rd.V(), rn.V());
}
@@ -1064,6 +1111,34 @@ public:
}
// Floating-point data-processing (2 source)
void fmul(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0000, rd, rn, rm);
}
void fdiv(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0001, rd, rn, rm);
}
void fadd(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0010, rd, rn, rm);
}
void fsub(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0011, rd, rn, rm);
}
void fmax(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0100, rd, rn, rm);
}
void fmin(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0101, rd, rn, rm);
}
void fmaxnm(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0110, rd, rn, rm);
}
void fminnm(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0111, rd, rn, rm);
}
void fnmul(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b1000, rd, rn, rm);
}
void fmul(SRegister rd, SRegister rn, SRegister rm) {
Float2Source(0, 0, 0b00, 0b0000, rd.V(), rn.V(), rm.V());
}
@@ -1149,6 +1224,16 @@ public:
}
// Floating-point conditional select
void fcsel(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, Condition Cond) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for {}", __func__);
const uint32_t ConvertedSize =
size == ScalarRegSize::i64Bit ? 0b01 :
size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
FloatConditionalSelect(0, 0, ConvertedSize, rd, rn, rm, Cond);
}
void fcsel(SRegister rd, SRegister rn, SRegister rm, Condition Cond) {
FloatConditionalSelect(0, 0, 0b00, rd.V(), rn.V(), rm.V(), Cond);
}
@@ -1304,6 +1389,16 @@ private:
dc32(Instr);
}
void Float1Source(ScalarRegSize size, uint32_t M, uint32_t S, uint32_t opcode, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for {}", __func__);
const uint32_t ConvertedSize =
size == ScalarRegSize::i64Bit ? 0b01 :
size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
Float1Source(M, S, ConvertedSize, opcode, rd, rn);
}
// Floating-point compare
void FloatCompare(uint32_t M, uint32_t S, uint32_t ftype, uint32_t op, uint32_t opcode2, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0001'1110'0010'0000'0010'0000'0000'0000;
@@ -1336,6 +1431,7 @@ private:
dc32(Instr);
}
// Floating-point data-processing (2 source)
void Float2Source(uint32_t M, uint32_t S, uint32_t ptype, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0001'1110'0010'0000'0000'1000'0000'0000;
@@ -1350,6 +1446,16 @@ private:
dc32(Instr);
}
void Float2Source(ScalarRegSize size, uint32_t M, uint32_t S, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for {}", __func__);
const uint32_t ConvertedSize =
size == ScalarRegSize::i64Bit ? 0b01 :
size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
Float2Source(M, S, ConvertedSize, opcode, rd, rn, rm);
}
// Floating-point conditional select
void FloatConditionalSelect(uint32_t M, uint32_t S, uint32_t ptype, VRegister rd, VRegister rn, VRegister rm, Condition Cond) {
uint32_t Instr = 0b0001'1110'0010'0000'0000'1100'0000'0000;
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
/* System instruction emitters.
*
* This is mostly a mashup of various instruction types.
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/BlockSamplingData.h"
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstdint>
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
@@ -16,6 +17,12 @@ constexpr static uint64_t NamedVectorConstants[FEXCore::IR::NamedVectorConstant:
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT_UPPER
{0x0000'0001'0000'0000ULL, 0x0000'0003'0000'0002ULL}, // NAMED_VECTOR_MOVMSKPS_SHIFT
{0x040B'0E01'0B0E'0104ULL, 0x0C03'0609'0306'090CULL}, // NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE
{0x0706'0504'FFFF'FFFFULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0110B
{0x0706'0504'0302'0100ULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0111B
{0xFFFF'FFFF'0302'0100ULL, 0x0F0E'0D0C'FFFF'FFFFULL}, // NAMED_VECTOR_BLENDPS_1001B
{0x0706'0504'0302'0100ULL, 0x0F0E'0D0C'FFFF'FFFFULL}, // NAMED_VECTOR_BLENDPS_1011B
{0xFFFF'FFFF'0302'0100ULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1101B
{0x0706'0504'FFFF'FFFFULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1110B
};
constexpr static auto PSHUFLW_LUT {
@@ -126,6 +133,145 @@ constexpr static auto PSHUFD_LUT {
}()
};
constexpr static auto SHUFPS_LUT {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// 32-bit words in [127:96], [95:64], [63:32], [31:0] are selected using the 8-bit Index.
// Expectation for this LUT is to simulate SHUFPS with ARM's TBL (two register) instruction.
// SHUFPS behaviour:
// Two 32-bits words from each source are selected from each source in the lower and upper halves of the 128-bit destination.
// Dest[31:0] = Src1[<Word0>]
// Dest[63:32] = Src1[<Word1>]
// Dest[95:64] = Src2[<Word2>]
// Dest[127:96] = Src2[<Word3>]
std::array<LUTType, 256> TotalLUT{};
const uint64_t WordSelectionSrc1[4] = {
0x03'02'01'00,
0x07'06'05'04,
0x0b'0a'09'08,
0x0f'0e'0d'0c,
};
// Src2 needs to offset each byte index by 16-bytes to pull from the second source.
const uint64_t WordSelectionSrc2[4] = {
0x03'02'01'00 + (0x10101010),
0x07'06'05'04 + (0x10101010),
0x0b'0a'09'08 + (0x10101010),
0x0f'0e'0d'0c + (0x10101010),
};
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] =
(WordSelectionSrc1[Word0] << 0) |
(WordSelectionSrc1[Word1] << 32);
LUT.Val[1] =
(WordSelectionSrc2[Word2] << 0) |
(WordSelectionSrc2[Word3] << 32);
}
return TotalLUT;
}()
};
constexpr static auto DPPS_MASK {
[]() consteval {
struct LUTType {
uint32_t Val[4];
};
std::array<LUTType, 16> TotalLUT{};
for (size_t i = 0; i < TotalLUT.size(); ++i) {
auto &LUT = TotalLUT[i];
constexpr auto GetLUT = [](size_t i, size_t Index) {
if (i & (1U << Index)) {
return -1U;
}
return 0U;
};
LUT.Val[0] = GetLUT(i, 0);
LUT.Val[1] = GetLUT(i, 1);
LUT.Val[2] = GetLUT(i, 2);
LUT.Val[3] = GetLUT(i, 3);
}
return TotalLUT;
}()
};
constexpr static auto DPPD_MASK {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
std::array<LUTType, 4> TotalLUT{};
for (size_t i = 0; i < TotalLUT.size(); ++i) {
auto &LUT = TotalLUT[i];
constexpr auto GetLUT = [](size_t i, size_t Index) {
if (i & (1U << Index)) {
return -1ULL;
}
return 0ULL;
};
LUT.Val[0] = GetLUT(i, 0);
LUT.Val[1] = GetLUT(i, 1);
}
return TotalLUT;
}()
};
constexpr static auto PBLENDW_LUT {
[]() consteval {
struct LUTType {
uint16_t Val[8];
};
// 16-bit words in [127:112], [111:96], [95:80], [79:64], [63:48], [47:32], [31:16], [15:0] are selected using 8-bit swizzle.
// Expectation for this LUT is to simulate PBLENDW with ARM's TBX (one register) instruction.
// PBLENDW behaviour:
// 16-bit words from the source is moved in to the destination based on the bit in the swizzle.
// Dest[15:0] = Swizzle[0] ? Src[15:0] : Dest[15:0]
// Dest[31:16] = Swizzle[1] ? Src[31:16] : Dest[31:16]
// Dest[47:32] = Swizzle[2] ? Src[47:32] : Dest[47:32]
// Dest[63:48] = Swizzle[3] ? Src[63:48] : Dest[63:48]
// Dest[79:64] = Swizzle[4] ? Src[79:64] : Dest[79:64]
// Dest[95:80] = Swizzle[5] ? Src[95:80] : Dest[95:80]
// Dest[111:96] = Swizzle[6] ? Src[111:96] : Dest[111:96]
// Dest[127:112] = Swizzle[7] ? Src[127:112] : Dest[127:112]
std::array<LUTType, 256> TotalLUT{};
const uint16_t WordSelectionSrc[8] = {
0x01'00,
0x03'02,
0x05'04,
0x07'06,
0x09'08,
0x0B'0A,
0x0D'0C,
0x0F'0E,
};
constexpr uint16_t OriginalDest = 0xFF'FF;
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
for (size_t j = 0; j < 8; ++j) {
LUT.Val[j] = ((i >> j) & 1) ? WordSelectionSrc[j] : OriginalDest;
}
}
return TotalLUT;
}()
};
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState), InitialCodeSize(InitialCodeSize), MaxCodeSize(MaxCodeSize) {
@@ -136,10 +282,17 @@ CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t I
Common.NamedVectorConstantPointers[i] = reinterpret_cast<uint64_t>(NamedVectorConstants[i]);
}
// Copy named vector constants.
memcpy(Common.NamedVectorConstants, NamedVectorConstants, sizeof(NamedVectorConstants));
// Initialize Indexed named vector constants.
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFLW] = reinterpret_cast<uint64_t>(PSHUFLW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFHW] = reinterpret_cast<uint64_t>(PSHUFHW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFD] = reinterpret_cast<uint64_t>(PSHUFD_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_SHUFPS] = reinterpret_cast<uint64_t>(SHUFPS_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPS_MASK] = reinterpret_cast<uint64_t>(DPPS_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPD_MASK] = reinterpret_cast<uint64_t>(DPPD_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PBLENDW] = reinterpret_cast<uint64_t>(PBLENDW_LUT.data());
#ifndef FEX_DISABLE_TELEMETRY
// Fill in telemetry values
Loaded 100 of 741 files, more files were not shown because too many files have changed in this diff. Show more