Compare commits

..
812 Commits
Author SHA1 Message Date
Ryan Houdek d1a4029bc5 Docs: Update for release FEX-2503 2025-03-05 09:50:20 -08:00
Ryan Houdek 97070aad25 Merge pull request #4381 from Sonicadvance1/fix_double_load
ArgumentLoader: Fixes double load
2025-03-05 09:45:46 -08:00
Ryan Houdek 39640185a3 Merge pull request #4383 from Sonicadvance1/remove_uninitialized_variables
Various: Removes warnings about uninitialized variables
2025-03-05 09:45:27 -08:00
LC 4b17506ffe Merge pull request #4382 from Sonicadvance1/fix_profiler_crash
Profiler: Fixes potential crash due to uninitialized variables
2025-03-04 23:53:32 -05:00
LC 2435ebecbe Merge pull request #4384 from Sonicadvance1/codeemitter_missing_checks
CodeEmitter: Adds missing assert checks
2025-03-04 23:52:19 -05:00
Ryan Houdek 410a35b968 unittests/CodeEmitter: Fixes incorrect test values 2025-03-04 20:29:03 -08:00
Ryan Houdek c8d234a767 CodeEmitter: Adds missing assert checks
We weren't checking if the post-index variants of these instructions
were using the correct post-offset. These support /only/ the correctly
sized post-index. The no-offset version is an entirely different set of
functions.
2025-03-04 20:18:16 -08:00
Ryan Houdek 2bf87ff40f Various: Removes warnings about uninitialized variables
NFC. These wouldn't even occur in practice.
2025-03-04 19:59:26 -08:00
Ryan Houdek ba347c49c9 Profiler: Fixes potential crash due to uninitialized variables
If ProfileStats aren't enabled then `Initialize` early returns, but some
of these values weren't being zero initialized which could result in
crashes.

Ensure all the values in StatAlloc are zero initialized so this doesn't
occur.
2025-03-04 17:37:28 -08:00
Ryan Houdek 81434cd233 Merge pull request #4377 from bylaws/sidt
Implement SIDT/LSL
2025-03-04 17:10:32 -08:00
Ryan Houdek 46d0df9cba ArgumentLoader: Fixes double load
The argument loader was loading configuration from the arguments twice.

The use of the argument loader needs to preload the arguments before
being handed off to the config system. This way we can pull remaining
arguments that get passed to the guest application.

Due to this, `Load` was getting called twice, once in the constructor
and once in the Config system. This was causing the backend to allocate
twice as much memory since the second load appends the arguments to a
fextl::list internally.

Not really any functional change but it was causing some heartburn with
some changes I was working on.
2025-03-04 16:51:56 -08:00
Ryan Houdek b7f58e68c5 Merge pull request #4363 from pmatos/ReciprocalsFix
Improve reciprocal estimate and tests
2025-03-01 00:58:45 -08:00
Billy Laws 6ba2accbdc Frontend: Mark INVLPG as permission-restricted 2025-02-28 16:22:47 +00:00
Billy Laws 29ee94d331 Frontend: Correct TYPE_SECOND_GROUP_MODRM handling 2025-02-28 16:02:24 +00:00
Billy Laws cdae654fe4 FEXCore: Somewhat implement LSL
Emulate by always returning failure, this deviates from both Linux
and Windows but shouldn't be depended on by anything.
2025-02-27 23:45:14 +00:00
Billy Laws a506c84bc7 FEXCore: Implement SIDT 2025-02-27 23:45:14 +00:00
Paulo Matos 8e4a47181b instcountci: Improve reciprocal estimate and tests 2025-02-27 15:14:39 +01:00
Paulo Matos b5241e0f60 Improve reciprocal estimate and tests
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.

* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.

Fixes #4319.
2025-02-27 15:07:14 +01:00
LC 57ed466a7f Merge pull request #4373 from Sonicadvance1/sha_data_shuffle_tbl
OpcodeDispatcher: Reuse PSHUFD shuffle mask for sha data shuffling
2025-02-25 11:22:53 -05:00
LC 4f46f55f2d Merge pull request #4374 from Sonicadvance1/float_packed_min_max_afp
JIT: Optimize packed float min/max if AFP is supported
2025-02-25 11:21:39 -05:00
Ryan Houdek 69cfc78ee1 InstcountCI: Update 2025-02-24 14:54:37 -08:00
Ryan Houdek 04f1ab8571 unittests: Extend minmax nan test for 64-bit
This was only testing 32-bit values before.
2025-02-24 14:54:13 -08:00
Ryan Houdek 95694b2017 JIT: Optimize packed float min/max if AFP is supported
If AFP.AH is supported then fmin/fmax behaves like the x86 min/max
instruction so we don't need to jump through any additional hoops.
Support this use case to save a few instructions when AFP is supported.
2025-02-24 14:53:18 -08:00
Ryan Houdek 00bed2f0c0 InstcountCI: Update 2025-02-24 12:08:12 -08:00
Ryan Houdek e718fc35f8 OpcodeDispatcher: Reuse PSHUFD shuffle mask for sha data shuffling
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.

OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).

Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
2025-02-24 12:07:48 -08:00
LC 717015bae8 Merge pull request #4364 from Sonicadvance1/pid_wine
FEXpidof: Fixes searching for wine applications
2025-02-23 15:22:16 -05:00
LC 530d3d809b Merge pull request #4371 from Sonicadvance1/update_vixl
Update vixl to ff82b3328c59fa4cf2fe36697b44eae15a650371
2025-02-23 10:01:25 -05:00
Ryan Houdek 596b32d15c Update simulator expectations 2025-02-23 03:31:43 -08:00
Ryan Houdek 8d3918b4f0 CodeEmitter: Update tests and new assert for unallocated instruction encoding 2025-02-23 03:31:43 -08:00
Ryan Houdek 4c7e31513b InstcountCI: Update 2025-02-23 02:53:17 -08:00
Ryan Houdek 765509d7f5 Update vixl 2025-02-23 02:52:57 -08:00
LC dbb58d10a6 Merge pull request #4370 from Sonicadvance1/sha1rnds4
OpcodeDispatcher: Emulate SHA1RNDS4 with ARM sha extensions
2025-02-22 14:28:33 -05:00
Ryan Houdek d448976b3c InstcountCI: Update 2025-02-22 04:57:59 -08:00
Ryan Houdek de6931b1f5 OpcodeDispatcher: Emulate SHA1RNDS4 with ARM sha extensions
```diff
     "sha1rnds4 xmm0, xmm1, 10b": {
-      "ExpectedInstructionCount": 55,
+      "ExpectedInstructionCount": 10,
```

So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
2025-02-22 04:57:58 -08:00
Ryan Houdek 35268d185e InstcountCI: Add sha1rnds4 to crypto file 2025-02-22 04:48:18 -08:00
LC beef9eee0a Merge pull request #4368 from Sonicadvance1/more_pshufd
OpcodeDispatcher: Implements a few more pshufd masks
2025-02-19 16:50:23 -05:00
LC 34a274d4e6 Merge pull request #4367 from Sonicadvance1/sha1_msg2
OpcodeDispatcher: Implement support for SHA1MSG2 using SHA instructions
2025-02-19 16:49:02 -05:00
Ryan Houdek ef28a6c19a InstcountCI: Update 2025-02-19 12:47:29 -08:00
Ryan Houdek 50b5971ee5 OpcodeDispatcher: Implements a few more pshufd masks
Saw these while scanning around. Funnily it makes it look like libnss is
worse off because there are multiple instructions using the same table
lookup to swizzle. So one instruction turns in to two.

We don't have a way to choose one path or the other, so it's usually
better to go the route that the instruction in a vacuum is improved, so
on average it is also improved.
2025-02-19 12:46:54 -08:00
Ryan Houdek 9a70ae18ea InstcountCI: Update 2025-02-19 11:35:56 -08:00
Ryan Houdek d10853b775 OpcodeDispatcher: Implement support for SHA1MSG2 using SHA instructions
Only saves a handful of instructions, but still an improvement.

```
   "sha1msg2 xmm0, xmm1": {
     -      "ExpectedInstructionCount": 11,
     +      "ExpectedInstructionCount": 7,
```
2025-02-19 11:33:52 -08:00
Ryan Houdek cdf6a16efc JIT: Implement ARM VSha1SU1 IR operation 2025-02-19 11:33:38 -08:00
LC 02d7261f51 Merge pull request #4366 from Sonicadvance1/sha256msg2_opt
OpcodeDispatcher: Implement SHA256MSG2 using SHA256 operation
2025-02-19 06:32:46 -05:00
Ryan Houdek 9c9ddeffbe InstcountCI: Update 2025-02-18 18:03:31 -08:00
Ryan Houdek afa8b3a5c9 OpcodeDispatcher: Implement SHA256MSG2 using new SHA256 operation 2025-02-18 18:03:31 -08:00
Ryan Houdek bbcd4c168c JIT: Implement support for VSha256U1 operation 2025-02-18 18:03:31 -08:00
LC d22bd9cac7 Merge pull request #4365 from Sonicadvance1/jit_code_tail_size
CPUBackend: Move bool to end of JITCodeTail
2025-02-18 17:21:42 -05:00
Ryan Houdek 5ccf25196e CPUBackend: Move bool to end of JITCodeTail
Reduces the size by 8 bytes from 48 to 40.
2025-02-18 11:45:30 -08:00
Ryan Houdek caf15a2dac Merge pull request #4359 from neobrain/feature_libfwd_fexconfig
Library Forwarding: Add GUI for enabling use of individual host libraries
2025-02-17 16:59:43 -08:00
Ryan Houdek 982a05450c FEXpidof: Fixes searching for wine applications
I kept finding I needed `./fex_shm_stats_read `FEXpidof Celeste.exe``
but FEXpidof wasn't ever wired up to find FEX in the face of emulating
wine and arm64 wine.

This adds two new features basically:
- If x86 wine is being emulated, then walk the argument list just like
  our config options to see what the program executable name is.
- If it is arm64 wine using FEX, then we need to detect that, and walk
  the arguments in a similar fashion

The detection is the main thing here in that the only way to detect FEX
for arm64 wine is checking the applications mapped files and seeing if
it is mapping arm64ecfex.dll or wow64fex.dll.

x86 Wine is easy since that's just skipping the wine{64,}{-preloader,}
arguments to get to the executable name.
2025-02-17 13:05:13 -08:00
Ryan Houdek 3e381b742c FHU: Support std::string_view GetFilename
The previous fextl::string version makes a copy. Theoretically most uses
of this function doesn't need a copy but there's a lot of dependencies
that would need to be converted for that.

So just add the string_view version.
2025-02-17 13:05:11 -08:00
Ryan Houdek 6b82664166 Merge pull request #4362 from neobrain/refactor_remove_unused
Remove unused code in various places
2025-02-16 18:31:42 -08:00
Tony Wasserka bb30a2eb1e CodeEmitter: Remove unused member function 2025-02-16 16:30:05 +01:00
Tony Wasserka b3fdf5c48f Core: Remove unused ThreadAddBlockLink interface 2025-02-16 16:30:05 +01:00
Tony Wasserka 17d5ed847f Core: Remove redundant lock_guard
LookupCache::Erase already acquires its mutex internally
2025-02-16 16:30:05 +01:00
Tony Wasserka 4335d17fc0 Core: Remove ContextImpl::AddBlockMapping interface
This was only used internally and doesn't add anything over using the
equivalent InternalThreadState interface directly.
2025-02-16 16:30:05 +01:00
Tony Wasserka ae07958577 Library Forwarding: Add GUI for enabling use of individual host libraries 2025-02-16 13:15:40 +01:00
Tony Wasserka c51b9ba3d6 Library Forwarding: Remove obsolete libraries from ThunksDB.json 2025-02-16 13:15:40 +01:00
Tony Wasserka cc6ff5e9e6 Config: When saving Config.json, preserve ThunksDB entries 2025-02-16 13:15:40 +01:00
Tony Wasserka b968ea7e7e Windows: Add atoll symbol used by json_getInteger 2025-02-16 13:15:40 +01:00
Ryan Houdek d14b6e160e Merge pull request #4360 from neobrain/fix_libfwd_build
Library Forwarding: Fix build problems on some platforms
2025-02-15 20:48:06 -08:00
LC b09b9488ef Merge pull request #4361 from neobrain/fix_tracy_log
Profiler: Drop accidentally included debugging code
2025-02-15 12:38:31 -05:00
Tony Wasserka a8120ee7ef Profiler: Drop accidentally included debugging code 2025-02-15 15:38:28 +01:00
Tony Wasserka fb82059750 Library Forwarding: Fix build on platforms that put headers for libwayland-client in a subfolder 2025-02-15 14:53:16 +01:00
Tony Wasserka 0fc6240d72 Library Forwarding/GL: Only export GLX entrypoints available at compile-time 2025-02-15 14:53:16 +01:00
Ryan Houdek a7c6fdb1fc Merge pull request #4357 from Sonicadvance1/fix_negative_return
FileManagement: Throw a warning if `/proc` can't be opened
2025-02-13 10:55:00 -08:00
Ryan Houdek b76a2963cf Merge pull request #4355 from Sonicadvance1/move_instead_of_copy
Fixes a couple locations where is variable is copied when it could be moved
2025-02-13 10:54:46 -08:00
Ryan Houdek 42b0fbd34c Merge pull request #4354 from Sonicadvance1/use_of_auto_copy
Fixes some instances of auto usage with unintentional copy
2025-02-13 10:54:29 -08:00
Ryan Houdek f69ef8606f Merge pull request #4353 from Sonicadvance1/seccomp_fixes
Seccomp: Fix a couple minor things.
2025-02-13 10:54:12 -08:00
Ryan Houdek afa5ad5f9f Merge pull request #4351 from Sonicadvance1/fix_pagesize_check
Linux: Fixes PAGESIZE checks that could return <= 0
2025-02-13 10:54:01 -08:00
Ryan Houdek 1fc82708e9 Merge pull request #4349 from Sonicadvance1/relative_portable
Config: Correctly handle relative paths with portable
2025-02-13 10:53:41 -08:00
Ryan Houdek 917cbbadde Fixes a couple locations where is variable is copied when it could be moved 2025-02-13 02:07:00 -08:00
LC c37dc81839 Merge pull request #4356 from Sonicadvance1/remove_elf_symbol_database
CommonTools: Removes ELFSymbolDatabase
2025-02-13 05:06:13 -05:00
LC df718d55ef Merge pull request #4358 from Sonicadvance1/arm64_unaligned
ArchHelpers/Arm64: Fix loadstore mask
2025-02-13 05:02:14 -05:00
LC 54412f1d5e Merge pull request #4352 from Sonicadvance1/profiler_fix_zero
Profiler: Fixes zeroing of allocated slots.
2025-02-13 04:54:31 -05:00
LC da76023bea Merge pull request #4350 from Sonicadvance1/fix_sve_fcvt
CodeEmitter: Minor fixes to SVE fcvtz{u,s}
2025-02-13 04:52:50 -05:00
Ryan Houdek 41e9309a36 ArchHelpers/Arm64: Fix loadstore mask
This would become an issue when multiple threads are contending with the SIGBUS handler on the same code.

We were failing to mask the VR, OPC, Rm, and Option bits, resulting in a
comparison below always resulting in a false result if another thread
managed to backpatch.

This was just unlikely to be seen on LRCPC2 supporting hardware and
since we fixed `LDSTUNSCALED_MASK` before, this wasn't really getting
seen.
2025-02-13 01:10:03 -08:00
Ryan Houdek 116268b275 FileManagement: Throw a warning if /proc can't be opened
We use this for ProcFD collision checking. Give a warning if it can't be
opened, also not doing the additional work when it fails.

This isn't likely to occur unless someone messes up their rootfs mounts.
2025-02-13 00:43:13 -08:00
Ryan Houdek 73802492b8 CommonTools: Removes ELFSymbolDatabase
This is completely unused.
2025-02-13 00:14:08 -08:00
Ryan Houdek 7cd4d53fa9 Fixes some instances of auto usage with unintentional copy
Just switch the uses over to `const auto&`
2025-02-12 23:51:57 -08:00
Ryan Houdek c44757975e Seccomp: Fix a couple minor things.
If fcntl fails then report a log message, and fix a potential overflow
before widen bug.
2025-02-12 23:45:06 -08:00
Ryan Houdek f25cdcdf63 Profiler: Fixes zeroing of allocated slots.
Was accidentally zeroing size of ThreadStatsHeader instead of
ThreadStats. So 64 bytes instead of 48, which would have overrunned a
slot.
2025-02-12 23:36:15 -08:00
Ryan Houdek 5d37253e85 Linux: Fixes PAGESIZE checks that could return <= 0
If any `sysconf(_SC_PAGESIZE);` errors then we can get bad values, make
sure to at minimum use the x86 page size.

Also changes a hardcoded page size to use the FEX pagesize define.
2025-02-12 23:32:15 -08:00
Ryan Houdek cd5f42ec79 CodeEmitter: Minor fixes to SVE fcvtz{u,s}
We had duplicated code paths in the ternary selection for 64-bit source
size, and on 32-bit source size 16-bit is an invalid target so the
second ternary was dead.
2025-02-12 23:09:41 -08:00
LC 6651f9e94b Merge pull request #4300 from neobrain/feature_profiler_tracy
Profiler: Add Tracy backend
2025-02-12 20:11:32 -05:00
Ryan Houdek e3ee579f92 Config: Correctly handle relative paths with portable
It is desired that FEX_APP_CONFIG and FEX_APP_CONFIG_LOCATION support
relative paths when portable is used. Support this.
2025-02-12 12:08:28 -08:00
Ryan Houdek 73e7240574 config_generator: Fix double FEX_ prefix for man options 2025-02-12 12:07:26 -08:00
Tony Wasserka 391f9aa97d Profiler: Add Tracy backend
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only

To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.

Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
2025-02-12 19:35:15 +01:00
Tony Wasserka 8ad54e7bd5 External: Add Tracy submodule 2025-02-12 19:25:35 +01:00
Ryan Houdek 9eccc01dd3 Merge pull request #4336 from Sonicadvance1/softfloat_stats
FEXCore/Profiler: Implement support for JIT float fallbacks
2025-02-11 17:31:56 -08:00
Ryan Houdek 39c1f816fc InstcountCI: Update 2025-02-11 14:41:24 -08:00
Ryan Houdek a32b892787 FEXCore/Profiler: Implement support for JIT float fallbacks
Based on #4291 and #4324. Ideally this gets merged at the same time so
we can have Mangohud be on version 2 before giving them an upstream
patch.

Performance-wise this change falls within noise of my x87 microbench.

This just lets us track the number of float fallbacks FEX does, letting
us detect things like x87 fallbacks and how frequent they are, so we can
detect if a game might be slow or stuttering because of these fallbacks.
2025-02-11 14:41:16 -08:00
Ryan Houdek 1b144ba3f0 Merge pull request #4347 from Sonicadvance1/pcmpistri_vector
FEXCore: Keep PCMPISTRI arguments in vectors longer
2025-02-11 14:29:12 -08:00
LC 6a39a8db72 Merge pull request #4291 from Sonicadvance1/profile_stats
FEX: Implements new sampling based stats
2025-02-11 16:51:35 -05:00
Ryan Houdek 602c530615 Wine: Add support for magic fex+wine shm path
Fallback to the previous path if it doesn't exist.
2025-02-11 13:42:45 -08:00
Ryan Houdek c8c27f26f7 Review 2025-02-11 13:42:45 -08:00
Ryan Houdek 549cdc4c2c InstcountCI: Update 2025-02-11 12:57:43 -08:00
Ryan Houdek 3160e0a430 FEXCore: Keep PCMPISTRI arguments in vectors longer
This reduces our codegen size and removes a few umov instructions.
Performance falls within noise but this small change will allow us to do
more vector optimizations in C code in the future.
2025-02-11 12:57:11 -08:00
Ryan Houdek dcebe85f3a Wine: Implements support for profile stats
This is a little trickier, we actually open the
`/dev/shm/fex-<pid>-stats` file directly using Windows APIs that way
Mangohud (which is going to be on the Linux side, or potentially even
embedded in to Gamescope) can safely pick up the stats.

A little quirky plus doesn't support expanding its size since WINE
doesn't support NtExtendSection, but that's fine.
2025-02-11 12:56:11 -08:00
Ryan Houdek 2ba0b66426 LinuxSyscalls: Implements support for Linux side profile stats
This is fairly straightforward. It creates the shared memory region in
/dev/shm/fex-<pid>-stats so that Mangohud can sample it.
2025-02-11 12:56:11 -08:00
Ryan Houdek 5c9543f159 Common: Implement a base profiler implementation
Not wired up to anything. Requires the frontends to allocate shared
memory in the expected way.
2025-02-11 12:56:11 -08:00
Ryan Houdek f6e3689f30 WinAPI: Implement support for DeleteFile 2025-02-11 12:56:11 -08:00
Ryan Houdek a761343717 Profiler: Sprinkle the profile stats around
For the four things we care about
2025-02-11 12:56:10 -08:00
Ryan Houdek b4c47a3d24 FEXCore: Implements baseline per-thread profile stats
Not wired up, just the definitions so it lives in the
InternalThreadState.

We want this accessible from both FEXCore and the frontends so it needs
to live there.

Two types of events supported. Scoped cyclecounts and instant
increments.

This gives us JIT time and Signal handling time, plus events for number
of SIGBUS and number of SMC events.

All useful statistics for seeing stutter live.
2025-02-11 12:56:10 -08:00
Ryan Houdek 906988c49b Windows: Expose support for NtCreateSection and NtMapViewOfSection 2025-02-11 12:56:10 -08:00
Ryan Houdek 4186b2ad82 Merge pull request #4346 from neobrain/fix_unused_header
Remove unused IMGui header and obsolete debugger documentation
2025-02-11 10:09:53 -08:00
LC 2943cff73f Merge pull request #4345 from Sonicadvance1/optimize_pcmpistri
FEXCore: Optimize VPCMPISTRX implicit length calculation
2025-02-11 12:50:55 -05:00
Tony Wasserka 53ac5579fb Remove unused IMGui header and obsolete debugger documentation 2025-02-11 15:41:35 +01:00
LC 0d7a9f911a Merge pull request #4324 from Sonicadvance1/vector_reg_x87
FEXCore/JIT: Pass Softfloat arguments as vector registers
2025-02-11 08:00:55 -05:00
LC dce9de222d Merge pull request #4344 from Sonicadvance1/fix_fexserver_compressedimage_start
FEXServer: Fixes background startup
2025-02-11 07:56:37 -05:00
Ryan Houdek d46722a95c FEXCore: Optimize VPCMPISTRX implicit length calculation
With ASIMD this can be decently faster. With my microbenchmark this
makes pcmpistri ~6% faster.

With #4324 this can be made even faster since the incoming data can stay
in vector registers; Removing some overhead of umov.
2025-02-10 20:17:29 -08:00
Ryan Houdek 3ba4da7736 InstcountCI: Update 2025-02-10 12:54:06 -08:00
Ryan Houdek 6abf5b90b7 FEXCore/JIT: Pass Softfloat arguments as vector registers
This is preparation work to allow passing the corestate to the x87 soft
float handlers directly for some profile stats.

Performance-wise, this change falls within noise because it basically
moves the GPR->Vector moves from the JIT in to C code, my microbench saw
the largest excursion of 5% but that's still within noise in the current
design of my bench.

A more tangible win from this change alone is less codegen on the JIT
side.
2025-02-10 12:53:25 -08:00
Ryan Houdek 00aa4ddea0 FEXCore/Softfloat: Support loading and storing SoftFloat to vector registers 2025-02-10 12:53:25 -08:00
Ryan Houdek d0c6f9de22 External/vixl: Update 2025-02-10 12:53:25 -08:00
Ryan Houdek a85cc85081 Merge pull request #4341 from Sonicadvance1/4216_#2
OpcodeDispatcher: Use offset for LRCPC2 more frequently
2025-02-10 11:37:23 -08:00
Ryan Houdek 75793300f2 Merge pull request #4333 from neobrain/feature_fasio
Async: Add framework for multiplexing IO on network sockets and other file descriptors
2025-02-10 11:35:37 -08:00
Ryan Houdek 672805584e InstcountCI: Update 2025-02-10 10:38:57 -08:00
Ryan Houdek 5b4fd590d1 OpcodeDispatcher: Use offset for LRCPC2 more frequently
We were missing small offset immediate encoded LRCPC2 pretty much
always.
This fixes that. Finishes up what #4216 started.
2025-02-10 10:38:39 -08:00
Ryan Houdek 0ccd38f593 JIT: Fixes offset for LRCPC2 LoadStoreMemTSO
This was in an assert statement which wouldn't give us the offset.
2025-02-10 10:38:39 -08:00
Alyssa Rosenzweig b46e5d4488 Merge pull request #4342 from Sonicadvance1/store_as_zero
JIT: Optimize memory stores with zero
2025-02-10 09:59:44 -05:00
Alyssa Rosenzweig d7223d598f Merge pull request #4340 from Sonicadvance1/4216_#1
InstCountCI: fix turnip instcountci
2025-02-10 09:58:30 -05:00
Ryan Houdek 7a0368132d Merge pull request #4343 from Sonicadvance1/fix_fexserver_search
FEXServerClient: Fix searching for FEXServer
2025-02-10 02:33:34 -08:00
Ryan Houdek aff3914a66 FEXServerClient: Fix searching for FEXServer
argv[0] is whatever the user passed in and may not directly be
FEXLoader/FEXInterpreter's path. Make sure get the full path.
2025-02-10 01:22:47 -08:00
Ryan Houdek 8876047875 FEXServer: Fixes background startup
The problem here is that the pipe we used for telling FEXInterpreter
that the FEXServer is ready to accept connections was inherited by
erofsfuse or squashfuse. So the closing of the pipe from the FEXServer
side would leave a reference open in squashfuse or erofsfuse.

Fix this by setting FD_CLOEXEC on the pipe, but also pass the pipe FD
through an argument instead of scanning for all pipes.

Then once we execve the squashfuse/erofsfuse application, the FD isn't
inherited.

Fixes #4329
2025-02-09 23:59:32 -08:00
Ryan Houdek 78e2aa16f0 InstcountCI: Update 2025-02-09 22:52:22 -08:00
Ryan Houdek 3ef695cf70 JIT: Optimize memory stores with zero
Minor optimization but I've seen it around.
2025-02-09 22:50:12 -08:00
Alyssa Rosenzweig c4d8dd6413 InstCountCI: fix turnip instcountci
this was 32-bit

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-02-08 21:41:16 -08:00
Ryan Houdek a49d30f6e2 Merge pull request #4327 from bylaws/mbdef
Config: Enable multiblock by default
2025-02-08 02:54:03 -08:00
Ryan Houdek 512643d3d6 Docs: Update for release FEX-2502 2025-02-08 01:04:19 -08:00
Ryan Houdek b3a69af752 Merge pull request #4338 from Sonicadvance1/fix_vl_int16
FEXCore/vl64: Fixes int16 encoding
2025-02-07 19:08:28 -08:00
Ryan Houdek 64c0dc47a9 unittests/FEXCore: Fixes VL test and adds decode check
So when encoding we also test the decode path.
2025-02-07 15:30:16 -08:00
Ryan Houdek 923c323d6f FEXCore/vl64: Fixes int16 encoding
For some reason when I was writing the tests I got the byte order
incorrect. The type header needs to be in the first byte, not the second
byte.
2025-02-07 15:29:51 -08:00
Ryan Houdek ee47b5bbc9 Merge pull request #4331 from Sonicadvance1/fix_portable_fexserver
FEXServer: Fixes FEX_PORTABLE usage
2025-02-07 12:15:58 -08:00
LC e8cd655c84 Merge pull request #4330 from Sonicadvance1/hotblock_tso_32bit
InstcountCI: Adds a hotblock for 32-bit TSO testing
2025-02-07 14:56:19 -05:00
Tony Wasserka d80daf2692 FEXServerClient: Migrate RequestPIDFDPacket to fasio
The other operations in this file are simple reads/writes, so they don't
need to be changed.
2025-02-07 16:08:18 +01:00
Tony Wasserka 0e5c9e8b06 Async: Add helper for fixed-length reads 2025-02-07 16:08:17 +01:00
Tony Wasserka 6bc4aed82c Async: Fix receiving FDs via tcp_socket 2025-02-07 16:08:16 +01:00
Tony Wasserka e69e1200f5 Async: Qualify system call wrappers with :: 2025-02-07 10:50:40 +01:00
Tony Wasserka 81e253b06b Async: Handle EINTR and EAGAIN 2025-02-07 10:49:45 +01:00
Ryan Houdek 9af52fb642 Merge pull request #4335 from Sonicadvance1/fix_portable_wine
WINE: Fixes FEX_PORTABLE usage
2025-02-06 17:49:59 -08:00
Ryan Houdek eaddd44d17 Merge pull request #4326 from bylaws/mbfast2
Frontend: Split blocks at jump target boundaries
2025-02-06 17:48:56 -08:00
Ryan Houdek 854e699589 WINE: Fixes FEX_PORTABLE usage
Completely didn't listen to FEX_PORTABLE. Necessary otherwise it can
read configs from some random locations when portable is enabled.
2025-02-06 15:36:27 -08:00
Tony Wasserka 43dcc84c07 Async: Move ownership of file descriptors out of poll_reactor 2025-02-06 22:35:29 +01:00
Tony Wasserka 6e9d5f00de Async: Check function signature for callbacks 2025-02-06 22:35:29 +01:00
Tony Wasserka a65ca9663f Async: Rename read_callbacks to callbacks 2025-02-06 22:35:29 +01:00
Tony Wasserka 02b767c0ea GdbServer: Migrate to fasio 2025-02-06 22:35:29 +01:00
Tony Wasserka 4b1c1d266d FEXServer: Migrate ProcessPipe to fasio 2025-02-06 22:35:29 +01:00
Tony Wasserka d9bf140971 FEXServer: Migrate Logger to fasio 2025-02-06 22:35:29 +01:00
Tony Wasserka 1402776ba6 FEXServer: Clean up socket path setup
The character counting logic isn't actually needed, since bind() doesn't
need the exact byte length of the input data. Dropping the manual bookkeeping
cleans up this code considerably.
2025-02-06 22:35:29 +01:00
Tony Wasserka e36fb47d98 FEXServer: Use a pipe to register new log clients instead of signaling across threads 2025-02-06 22:35:29 +01:00
Tony Wasserka 2bb37357c0 FEXConfig: Migrate inotify monitoring to fasio 2025-02-06 22:35:29 +01:00
Tony Wasserka 7494ac7615 Add framework for multiplexing IO on network sockets and other file descriptors
The design leans heavily on Boost.Asio, a battle-tested library that's widely
used and that forms the basis of upcoming C++ networking support.
2025-02-06 22:35:29 +01:00
Tony Wasserka 44bc3fb90b fextl: Add std::move_only_function replacement 2025-02-06 22:30:45 +01:00
Ryan Houdek 20b00ecc9b Merge pull request #4332 from bylaws/ecdmsk
ARM64EC: Set EC_ENTRY_CPUAREA_REG at inline SMC dispatcher entry
2025-02-05 17:06:32 -08:00
Billy Laws 40662f947f OpcodeDispatcher: Only set mark _Break as setting RIP in the non-MB case
If we're starting a new block here then the newly started block won't have
set RIP and it is erroneous to set it.
2025-02-06 00:01:40 +00:00
Billy Laws 6e01934edc Frontend: Zero InstructionSize before decoding
Required for PeekByte to work correctly before decoding.
2025-02-06 00:01:17 +00:00
Billy Laws 151fc5e97f Frontend: Split blocks at jump target boundaries
With the prior approach, backwards jumps into existing blocks would
explore the overlapping part rather than splitting the block, generating
needless code and wasting time decoding. Similarly, the current block
wouldn't be split when it is extended to overlap with a pending jump target.

Solve this by tracking blocks in a sorted vector and splitting existing blocks
on jumps when appropriate, in order to avoid any possibility of overlapping
blocks, which would break the lookup, misaligned and zero instruction blocks
are disallowed.
2025-02-06 00:01:17 +00:00
Billy Laws 5f431dc776 Frontend: Track the current instruction start address 2025-02-06 00:01:17 +00:00
Billy Laws 5ec4d3125a ARM64EC: Set EC_ENTRY_CPUAREA_REG at inline SMC dispatcher entry 2025-02-06 00:00:38 +00:00
Ryan Houdek 8e511e7db4 FEXServer: Fixes FEX_PORTABLE usage
This was causing FEXServer to look in to global installed paths and
local paths for things when FEXServer was started.

Ensure it listens to FEX_PORTABLE so this doesn't occur.
This also requires us to scan both data directories and config
directories to find them.
2025-02-05 15:59:29 -08:00
Ryan Houdek 99b8046f03 InstcountCI: Adds a hotblock for 32-bit TSO testing 2025-02-05 13:44:51 -08:00
Ryan Houdek d39dea1ae3 Merge pull request #4325 from bylaws/ircopy2
FEXCore: Don't copy IR after compilation
2025-02-05 12:23:33 -08:00
LC e94643d5ca Merge pull request #4302 from Sonicadvance1/vl_jit_reconstruction
FEXCore/JIT: Encode the JITRIPReconstructionEntries using variable length integer
2025-02-05 14:19:15 -05:00
Ryan Houdek 1becbab0dc Merge pull request #4328 from Sonicadvance1/clang_thunks_default
CMake: Default enable clang thunk building
2025-02-04 12:00:53 -08:00
Ryan Houdek d49efb451e CMake: Default enable clang thunk building
We already mandate clang for building FEX and building thunks with clang
has been well tested since the PPA builder has been using it for a long
time.
2025-02-04 11:27:40 -08:00
Ryan Houdek d2a56ebd8c Merge pull request #4223 from Sonicadvance1/netstream_timeout
GdbServer: Implement new netstream that can be interrupted
2025-02-04 10:58:11 -08:00
Billy Laws 0e11a9b7ac FEXCore: Don't copy IR after compilation
This wastes a significant amount of time when the copies IR is promptly
thrown away after compilation anyway.
2025-02-04 16:52:16 +00:00
Billy Laws 68c77d12d4 FEXCore: Don't delete IRListView move constructor 2025-02-04 16:51:59 +00:00
Billy Laws 423e29ba42 Config: Enable multiblock by default 2025-02-04 16:40:02 +00:00
Ryan Houdek dcfbc2f20e FEXCore/JIT: Encode the JITRIPReconstructionEntries using variable length integer
When #2722 implemented this initially and #4271 switched over to signed
int16_t there was assumptions made that int16_t was a reasonable
trade-off in encoding size versus needing to deal with 8-bit values
being too small in some cases.

In the common case we are almost always encoding 8-bit values because
instructions are typically linear (and less than 15-bytes in size), but
16-bit was chosen because optimizing JIT and multiple instructions that
don't cause exceptions can add up to larger than 8-bit.

Instead of hardcoding 16-bit values, implement a variable length integer
class where ~96.8% of values are 8-bit encoded, and the remaining 3.19% are encoded using 16-bit.
Due to some constraints that #4271 put in place, we can basically
guarantee currently that branch targets are within 16-bit. The VL class
does support 32-bit and 64-bit as well so if we change behaviour then
nothing needs to change.

Some stats when running Sonic Mania with multiblock enabled.
Encoded integers: 3,504,907
Encoded 8-bit:    3,393,095 (96.8%)
Encoded 16-bit:     111,812 (3.19%)
Encoded 32/64-bit:        0

Encoded Size:       3,615,181 bytes (3.44MiB)
Fixed encoded size: 7,007,604 bytes (6.68MiB)

Definitely worth using and saves the headache of large RIP/PC offsets
causing problems.
2025-02-03 11:54:52 -08:00
Ryan Houdek 0b29c99fed review 2025-02-03 11:54:12 -08:00
Ryan Houdek 18556a9f75 Netstream: Use a std::variant 2025-02-03 11:54:12 -08:00
Ryan Houdek 02d93782ba GdbServer: Implement new netstream that can be interrupted
A major limitation of iostream is that you can't have reads or writes
with a safe interrupt. Instead rewrite the interface with Linux ppoll so
that these can be safely interrupted with a signal and return early.
2025-02-03 11:54:12 -08:00
Ryan Houdek 62cfc26262 Merge pull request #4322 from Sonicadvance1/protect_first_page_altstack
SignalDelegator: Protect first page of the altstack
2025-02-03 11:53:45 -08:00
Ryan Houdek b01a6b94e7 SignalDelegator: Protect first page of the altstack
When the alt-stack gets overflown then it is hard to see what went wrong
since the TLS variable is no longer accessible.

Protect the first page that contains the TLS variable.

Fixes #4320
2025-02-02 23:21:03 -08:00
Ryan Houdek 713ebf1476 Merge pull request #4315 from neobrain/refactor_irdumper_const
IRDumper: Allow const RA data
2025-01-31 14:57:05 -08:00
Ryan Houdek 26e50efdb2 Merge pull request #4317 from neobrain/change_fexserver_close_timeout
FEXServer: Lower close timeout
2025-01-31 14:56:49 -08:00
Ryan Houdek c8928999bf Merge pull request #4316 from neobrain/fix_check_catch2_version
CMake: Check for compatible Catch2 versions
2025-01-31 14:56:36 -08:00
Ryan Houdek e1f378c6cf Merge pull request #4318 from neobrain/fix_allocator_format_string
Allocator: Fix format string
2025-01-31 14:56:22 -08:00
Tony Wasserka d10222329c Allocator: Fix format string 2025-01-31 15:27:44 +01:00
Tony Wasserka 176fa7ab1d Lower FEXServer close timeout 2025-01-31 15:24:47 +01:00
Tony Wasserka ebb7137839 IRDumper: Allow const RA data 2025-01-31 15:20:23 +01:00
Tony Wasserka d7092a1231 CMake: Check for compatible Catch2 versions 2025-01-31 15:15:35 +01:00
Ryan Houdek 55cbb0b340 Merge pull request #4313 from pmatos/StartupSleepname
Add option StartupSleepProcName
2025-01-30 07:44:46 -08:00
Ryan Houdek db7fb56e9d Merge pull request #4311 from pmatos/RevertPredRA
Revert "Enable RA of SVE Predicate Registers"
2025-01-30 07:41:51 -08:00
Ryan Houdek 2bed7440a8 Merge pull request #4309 from pmatos/3DnowSkip
Skip 3DNow tests with precision issues
2025-01-30 07:40:18 -08:00
Ryan Houdek 0019bdecef Merge pull request #4314 from neobrain/fix_tso_ldr_bitmask
Arm64: Fix bitmask used to match load/store instructions
2025-01-30 06:32:28 -08:00
Tony Wasserka 51d355da30 Arm64: Fix bitmask used to match load/store instructions
When multiple threads simultaneously SIGBUS on the same address, one of them
will perform the backpatching while the other will detect the backpatched
instruction sequence and hence report the SIGBUS as "handled".

This typo broke the instruction detection logic: The second thread would
assume the source of the SIGBUS was unrelated to TSO emulation and hence
report the signal as unhandled (generally triggering program abortion).

In practice, this problem did not manifest as FEX does not currently share
CodeBuffers between threads.
2025-01-30 14:56:17 +01:00
Paulo Matos 28170fd723 Add option StartupSleepProcName
Sleeps only if current process matches this name. Leave empty to sleep
StartupSleep seconds on all processes.
2025-01-30 10:16:35 +01:00
Paulo Matos 44c65c35c8 Revert "Enable RA of SVE Predicate Registers"
This reverts commit fcbf0de05a.

The initial user of this code has been re-implemented in  b148cc6c.
This is not needed any longer so we're removing it.
2025-01-29 11:56:19 +01:00
Paulo Matos d8f8daf48a Skip 3DNow tests with precision issues
Fixes #4280
2025-01-29 08:37:55 +01:00
Ryan Houdek b148cc6ca3 Merge pull request #4292 from pmatos/EnsurePredCacheReset2
Predicate cache alternative implementation
2025-01-28 18:41:52 -08:00
Ryan Houdek 2a4c169fff Merge pull request #4306 from Sonicadvance1/robust_zero_length_envp
FEXLoader/ELFCodeLoader: Be robust against zero length environment variables
2025-01-28 18:41:33 -08:00
Ryan Houdek c2c84e4bd8 FEXLoader/ELFCodeLoader: Be robust against zero length environment variables
For some reason steamwebhelper is setting a zero length environment
variable. This was causing an assert to be raised early as the web
helper was starting up.

Just stop trying to memcpy the zero length string, gets steamwebhelper
working in the steam beta client again
2025-01-28 16:52:55 -08:00
LC c2f8b5b1ba Merge pull request #4299 from Sonicadvance1/fix_48bit_wine
FEX: Allocate a VMA allocator when running on a 48-bit VA
2025-01-28 19:26:31 -05:00
Ryan Houdek 3a33f554a0 FEXCore/unittests: Adds a FlexBitSet test
To ensure correctness
2025-01-28 16:07:51 -08:00
Ryan Houdek 11ce97655b Allocator: Still need to return memory regions to frontend 2025-01-28 16:07:51 -08:00
Ryan Houdek dc866538d4 FEXCore/Allocator: Ensure small reservations aren't used
Anything less than three pages can't be used for FEX allocations due to
VMA implementation details. Plus we may have reduced a single page
reservation to zero with the prior ObjectAlloc size reservation.
2025-01-28 16:07:51 -08:00
Ryan Houdek 48ad9e9a87 Allocator/FlexBitSet: Fixes a rounding issue with small allocation regions
When small regions were being used for VMA allocations (less than 64
pages), this function was truncating the result to zero. Resulting in
incorrect `LiveVMARegion` size calculations. It would calculate that the
FlexBitSet consumes zero bits of space, even though it needs to use at
least 2, or 3 if we actually want to allocate anything from that
LiveVMARegion.

This was noticed in this PR because our VMA region tracking is being
used more, which has a more likely chance to have small VMA regions for
allocating from. Cause a 1page allocation to try and use a 3 page VMA
region for allocation, but failing because the FlexBitSet size wasn't
calculated correctly.

- Page layout:
- [0x0, 0x1000):    struct LiveVMARegion
- [0x1000, 0x2000): FlexBitSet<uint64_t> UsedPages
-  ^ This space wasn't allocated/mprotected due to the size not
   calculating correctly.
- [0x2000, 0x3000): Memory for allocation
2025-01-28 16:07:51 -08:00
Ryan Houdek c75778abeb FEX: Allocate a VMA allocator when running on a 48-bit VA
When running on a system with a 48-bit VA, if FEX does any allocations
between us reserving the upper 128TB and the application running, then
/technically/ we are intersecting with the application's memory region
in the lower 47-bits.

This didn't typically result in any problems due to how ASLR works, but
if we did any large allocations (like #4291 wants with 128MB VMA region)
then these typically get pushed higher in the VA space.

Again not usually a problem, but if you happen to be running an
application that is using MAP_FIXED with hardcoded addresses then this
can stomp over FEX-Emu memory causing problems.

This is what happens with Wine, it reserves the upper-32MB of its 47-bit
VA space, which is /highly/ likely to stomp on FEX memory. In-fact it
likely occurs all the time, we just got lucky with whatever it was
clobbering wasn't used at the time.

On 39-bit VA systems this isn't a problem because the mmap fails
outright with a warning message from WINE.

Because we are already reserving the upper 128TB of VA space, instead
just always enable our allocator and use the regions that were reserved.
We need to be a little bit careful to ensure we don't accidentally
allocate more memory post-reservation but that just requires a small
adjustment to our unique_ptr and constructor for the 64BitAllocator.

This means /all/ FEX-Emu allocations will be in the upper 128TB VA space
when running 64-bit applications on a 48-bit VA system. Which is kind of
nice.

Fixes WINE in #4291 when the allocator stats are bumped to 128MB per
process.
2025-01-28 16:07:51 -08:00
LC fb2a59a67f Merge pull request #4308 from Sonicadvance1/gpuvis_stack_mem
Profiler/GPUViz: Stop allocating memory
2025-01-28 18:11:33 -05:00
LC f4c92756fc Merge pull request #4307 from Sonicadvance1/fix_4296
InstCountCI: Hardcode xchg instructions
2025-01-28 00:02:46 -05:00
Ryan Houdek 657c27556c Profiler/GPUViz: Stop allocating memory
These parsing strings are tiny, less than 64 bytes all the time. Just
stack allocate the buffer. Makes it safer to use during extenuating
circumstances as well, like SIGBUS and SIGSEGV.
2025-01-27 18:49:15 -08:00
Ryan Houdek 42e68d8544 InstCountCI: Hardcode xchg instructions
Nasm between the versions of 2.16.03 and 2.15.05 starting changing the
operand order of the instruction on us. Hard code both operand encodings
to ensure coverage.

Fixes #4296
2025-01-27 18:29:58 -08:00
Ryan Houdek ae69c4d895 Merge pull request #4305 from Sonicadvance1/fix_gpuviz_typo
Profiler/GPUViz: Fixes typo in instant TraceObject
2025-01-27 14:15:09 -08:00
Ryan Houdek 5a0db4d812 Profiler/GPUViz: Fixes typo in instant TraceObject
String parsing was adding a newline but then we failed to use it. I
don't think it caused any issues considering how infrequent instant
profiler objects are used.
2025-01-27 12:59:21 -08:00
Paulo Matos ddd241fe39 instcount: Ensure predicate cache is reset when control flow leaves block 2025-01-27 20:12:48 +01:00
Paulo Matos 3dc7b8d90a asm_tests: Ensure predicate cache is reset when control flow leaves block 2025-01-27 20:12:48 +01:00
Paulo Matos 0bccb1ece5 Ensure predicate cache is reset when control flow leaves block
Whenever the control float leaves the block, it might clobber the
predicate register so we reset the cache whenever that happens.

Fixes #4264
2025-01-27 20:12:43 +01:00
LC bf1e319d90 Merge pull request #4303 from Sonicadvance1/missing_clang_format
CodeEmitter: Fixes clang_format
2025-01-27 01:08:18 -05:00
Ryan Houdek bc6ae7feb4 CodeEmitter: Fixes clang_format 2025-01-26 19:05:50 -08:00
Ryan Houdek 8d6a43d708 Merge pull request #4290 from Sonicadvance1/fix_v6.13
LinuxEmulation: Ensure syscall wrapper declaration has CpuStateFrame as the first argument
2025-01-23 13:55:42 -08:00
Ryan Houdek bd1bca2c3a Merge pull request #4298 from neobrain/fix_libfwd_wl_regression
Library Forwarding/wayland: Fix regression caused by erroneous format
2025-01-23 13:43:18 -08:00
Ryan Houdek 9858ab7388 LinuxEmulation: Ensure syscall wrapper declaration has CpuStateFrame as the first argument
Otherwise crashes occur.
2025-01-23 12:26:10 -08:00
Tony Wasserka 1f6b69573c CI fix 2025-01-23 19:26:00 +01:00
Tony Wasserka 1c8c5b77f1 Library Forwarding/wayland: Fix regression caused by erroneous format
Auto-formatting turned this into "libwayland - client", making FEX fail to
load the host-side equivalent of this library.
2025-01-23 19:20:01 +01:00
Ryan Houdek e9bd037cf9 Merge pull request #4297 from pmatos/upload-art
Update upload-artifact action to v4
2025-01-23 09:22:28 -08:00
Paulo Matos 6f8353ab28 Update upload-artifact action to v4 2025-01-23 16:47:15 +01:00
LC c25720429d Merge pull request #4294 from neobrain/refactor_codeemitter_cleanups
CodeEmitter: Various cleanups
2025-01-23 01:47:32 -05:00
Tony Wasserka 276e9aded3 Merge pull request #4295 from neobrain/fix_changelog_script
Scripts: Fix indentation of changelog items
2025-01-22 13:40:08 -05:00
Tony Wasserka b88ac3359d Scripts: Fix indentation of changelog items
GitHub's markdown parser requires at least 2 spaces to open a new level
of indentation.
2025-01-22 19:24:54 +01:00
Tony Wasserka 264f3be8b4 CodeEmitter: Remove unused Bind validation logic 2025-01-22 18:25:35 +01:00
Tony Wasserka 4282f96d35 CodeEmitter: Use inline constexpr constants over constexpr functions 2025-01-22 18:25:35 +01:00
Tony Wasserka 9cdd759fc1 CodeEmitter: Convert template specialization into function overload 2025-01-22 18:25:35 +01:00
Tony Wasserka 403e8f8702 CodeEmitter: Unify SingleUseForwardLabel and ForwardLabel 2025-01-22 18:25:35 +01:00
Ryan Houdek 2e989e4262 Merge pull request #4293 from pmatos/CleanupCode
NFC: Code cleanup
2025-01-22 09:14:52 -08:00
Paulo Matos 5666a352d4 NFC: Code cleanup
Removing unused declarations.
Cleaning up unused headers and empty lines.
Avoiding static analysis warnings on `const auto` defaulting to int.
2025-01-22 10:22:25 +01:00
Ryan Houdek 1aa8c6f996 Merge pull request #4284 from neobrain/refactor_autoformat_inl
CodeEmitter: Auto-format .inl headers
2025-01-21 13:24:53 -08:00
Tony Wasserka 9882f53613 Scripts: Add inl files to reformat.sh 2025-01-21 21:28:57 +01:00
Tony Wasserka 8760c593ec CodeEmitter: Reformat inl files 2025-01-21 21:28:33 +01:00
Tony Wasserka ad695bdd59 CodeEmitter: Allow inl headers to be processed by external tooling 2025-01-21 21:28:33 +01:00
Ryan Houdek adff4bb1d7 Merge pull request #4289 from pmatos/PassThroughFPRs
Pass through FPRs argument
2025-01-21 09:23:09 -08:00
Ryan Houdek 42c931cf22 Merge pull request #4288 from OFFTKP/sext
Fix slight inaccuracy in test 3_F7_05_2
2025-01-21 09:22:14 -08:00
Paulo Matos 5d44dea47c Pass through FPRs argument 2025-01-21 18:06:41 +01:00
Ryan Houdek 56c95e3b36 Merge pull request #4285 from neobrain/fix_ptso_offsets
Fix crashes in Paranoid TSO mode
2025-01-21 08:31:11 -08:00
Ryan Houdek 840f306a7d Merge pull request #4287 from neobrain/refactor_warn_fixes
Fix warnings about unused objects
2025-01-21 08:30:45 -08:00
Ryan Houdek 9def89d5f8 Merge pull request #4286 from neobrain/refactor_dont_assume
Drop assume-asserting logging macros
2025-01-21 08:30:22 -08:00
offtkp 84c2f93dab Sign extend into RDX 2025-01-21 15:31:34 +02:00
Tony Wasserka b30733e2a7 Fix warnings about unused objects 2025-01-21 12:28:21 +01:00
Tony Wasserka da58e6a597 Fix warnings about unused variables 2025-01-21 12:07:33 +01:00
Tony Wasserka 229e7c5b61 LogManager: Remove assuming assert macros
Placing optimization hints everywhere interferes with debugging of
RelWithDebInfo builds, since the debugger won't be able to reliably
inspect variables or control flow. These hints are better placed on an
individual basis after identifying bottlenecks in a profiler.
2025-01-21 12:01:33 +01:00
Tony Wasserka 26685143be Update code formatting for logging macros 2025-01-21 12:01:33 +01:00
Tony Wasserka e54b9237c6 Drop use of assume-asserting logging macros 2025-01-21 12:01:33 +01:00
LC ac1b6d9482 Merge pull request #4283 from Sonicadvance1/v6.13_syscalls
LinuxSyscalls: Update for new v6.13 syscalls
2025-01-20 20:29:14 -05:00
LC bb6e98a6fc Merge pull request #4282 from Sonicadvance1/v6.13_drm
IoctlEmulation/drm: Update for v6.13
2025-01-20 20:28:58 -05:00
Tony Wasserka 32c75f06b3 Arm64: Drop unnecessary nops in memcpy/memset 2025-01-20 17:58:54 +01:00
Tony Wasserka a5de2d1008 Context: More broadly enable TSO emulation in paranoid TSO mode
Previously, many games would fail to run due to accidentally disabling
TSO emulation in most instructions.
2025-01-20 17:58:54 +01:00
Tony Wasserka f841912c75 Arm64: Implement indirect memory addressing in paranoid TSO mode 2025-01-20 17:58:54 +01:00
Tony Wasserka 3b8c36882d Merge pull request #4270 from bylaws/crosspg
Frontend: Disallow cross-page branches in multiblock
2025-01-20 07:17:36 -05:00
Ryan Houdek fca4c7e6bf LinuxSyscalls: Update for new v6.13 syscalls
Just four new *at variants of the xattr syscalls.
This will also let us use the *at variants for the non-at versions but I
didn't implement that optimization because this is brand new.
2025-01-19 18:41:30 -08:00
Ryan Houdek 5ffc611d13 IoctlEmulation/drm: Update for v6.13 2025-01-19 17:51:49 -08:00
Ryan Houdek 130f02647b Externals/drm: Update to v6.13 2025-01-19 17:49:29 -08:00
Ryan Houdek 981eea6ade Merge pull request #4271 from bylaws/soff
CPUBackend: Make guest RIP reconstruction offsets signed
2025-01-17 14:35:25 -08:00
Ryan Houdek 3f788eb4a8 Merge pull request #4266 from pmatos/FSTOpt
x87 fst/fld optimization for different addrmodes
2025-01-17 13:52:56 -08:00
Billy Laws 486dc974c4 CPUBackend: Make guest RIP reconstruction offsets signed
With multiblock enabled, host code generated from guest code with a
lower address may be placed after host code generated from guest code
with a higher address in a multiblock. As each guest RIP reconstruction
entry is always relative to the one before it the offset needs to be
signed to allow this.
2025-01-17 21:49:43 +00:00
Billy Laws 3b1fbbc766 Frontend: Disallow cross-page branches in multiblock
This avoids both the generation of multiblocks that cover massive spans
of guest code, which causes issues for both context reconstruction
overflowing the RIP offset and attempting to decode branch targets
in unmapped memory regions.

Once support for querying mappings from the FEX frontend is in place this
limit could be increased if necessary, but this seems fine for now.
2025-01-17 21:41:58 +00:00
Ryan Houdek d01db8f293 Merge pull request #4278 from neobrain/refactor_reduce_vixl_options
CMake: Simplify vixl-related options
2025-01-15 14:14:24 -08:00
LC fd09ded049 Merge pull request #4277 from bylaws/wine
Windows: Fix wine check
2025-01-15 13:55:49 -05:00
LC 8e2b4a306d Merge pull request #4260 from Sonicadvance1/profile_win32
Profiler: Setup for usage on Windows
2025-01-15 13:54:54 -05:00
Tony Wasserka 1d58f38aa5 CMake: Clarify that ENABLE_VIXL_SIMULATOR won't work in production 2025-01-15 17:08:59 +01:00
Tony Wasserka 5ff9a83b07 CMake: Drop COMPILE_VIXL_DISASSEMBLER option 2025-01-15 17:04:54 +01:00
Billy Laws f5decb5f83 Windows: Fix overcommit size logic in the wine path 2025-01-14 20:27:10 +00:00
Billy Laws 11fc49a0f8 Windows: Fix wine check
This did not work before :)
2025-01-14 20:27:03 +00:00
Alyssa Rosenzweig 48c03d747a Merge pull request #4273 from bylaws/earlyend
Frontend: End multiblocks early after hitting 2 consecutive null bytes
2025-01-14 12:34:25 -05:00
Billy Laws 643750817a Frontend: End multiblocks early after hitting 2 consecutive null bytes
'add [rax], al' is almost never seen in actual code so the assumption
can be made that we are most likely trying to explore garbage code and
that this will never be hit. If it is then code will be generated at
that point (where Entrypoint == true).
2025-01-14 17:12:11 +00:00
LC a52dd71e44 Merge pull request #4276 from neobrain/fix_vixl_tests
CMake: Compile vixl if ENABLE_VIXL_DISASSEMBLER is set
2025-01-14 11:54:11 -05:00
Alyssa Rosenzweig 8c02bd43df Merge pull request #4269 from bylaws/jumpext
JIT: Avoid OOB EC bitmap checks in ExitFunction
2025-01-14 11:47:58 -05:00
Alyssa Rosenzweig f635a12129 Merge pull request #4272 from bylaws/declimit
Frontend: Stop all decoding once MaxInst/DecodeBufferSize is reached
2025-01-14 11:40:23 -05:00
Paulo Matos 8191c4905b instcountci: x87 fst/fld optimization for different addrmodes 2025-01-14 16:20:51 +01:00
Paulo Matos 58a034b79d asm_tests: x87 fst/fld optimization for different addrmodes 2025-01-14 16:20:47 +01:00
Paulo Matos 2d53867668 x87 fst/fld optimization for different addrmodes
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.

Discussed in #4252.
2025-01-14 16:20:33 +01:00
Tony Wasserka 8a57fc5838 CMake: Compile vixl if ENABLE_VIXL_DISASSEMBLER is set 2025-01-14 13:01:49 +01:00
Billy Laws 5481e6d79a Frontend: Stop all decoding once MaxInst/DecodeBufferSize is reached
Currently FinalInstruction causes only to the currently decoding block
to be terminated, but that is not enough as both MaxInst and
DefaultDecodedBufferSize are global limits that apply across all blocks
within a multiblock.
2025-01-12 21:27:53 +00:00
Billy Laws c852a58ee3 JIT: Avoid OOB EC bitmap checks in ExitFunction 2025-01-12 21:25:50 +00:00
Ryan Houdek 8cfc016b3f Merge pull request #4265 from pmatos/RevertPredCache
Revert pred cache
2025-01-10 12:25:38 -08:00
Ryan Houdek 8c94b782c6 Merge pull request #4263 from pmatos/patch-1
Print arg type f80Bit
2025-01-10 09:17:49 -08:00
Paulo Matos 1dce4919f2 instcountci: Revert "Cache predicate register generation from pattern" 2025-01-10 12:53:15 +01:00
Paulo Matos cbda688e29 Revert "Cache predicate register generation from pattern"
This reverts commit 72a4063651.

Caused #4264
2025-01-10 12:52:11 +01:00
Paulo Matos 159ed07e68 Print arg type f80Bit 2025-01-10 09:14:49 +01:00
LC a18b2d0e17 Merge pull request #4262 from Sonicadvance1/fix_fileleak
Windows/CRT: Fixes FD leak
2025-01-09 19:51:25 -05:00
Ryan Houdek 4c9adab58d Windows/CRT: Fixes FD leak
Noticed that the FEXCore config file was open forever.
2025-01-09 15:27:12 -08:00
Ryan Houdek 4c9f1b105d CRT/IO: Fixes sharing rules when writing is used
Fixes trace file opening since it needs to share with other users
opening the file for writing.
2025-01-09 14:25:48 -08:00
Ryan Houdek 2290353295 Wine: Ensure the profiler is initialized. 2025-01-09 14:25:48 -08:00
Ryan Houdek c16bf09310 Profiler: Setup for usage on Windows
This will get gpuviz working under Wine.
2025-01-09 14:25:48 -08:00
LC 90db9486ce Merge pull request #4259 from Sonicadvance1/fix_4121
FEXConfig: Fixes instcount not being editable by keyboard
2025-01-08 23:12:50 -05:00
Ryan Houdek b79faa6207 FEXConfig: Fixes instcount not being editable by keyboard
Fixes #4121
2025-01-08 16:43:34 -08:00
LC 2293d3067a Merge pull request #4258 from Sonicadvance1/libraries
cmake: Adds some missing STATIC qualifiers
2025-01-07 19:52:17 -05:00
LC de431f113e Merge pull request #4257 from Sonicadvance1/remove_dup_n2
CPUID: Remove duplicated ARM Neoverse-N2
2025-01-07 19:51:29 -05:00
Ryan Houdek 34e265a801 cmake: Adds some missing STATIC qualifiers
Noticed this as I was scrolling through some cmake. Usually this doesn't
matter as we declare `BUILD_SHARED_LIBS` as False/Off, but this can
technically be overridden even when we don't want to.

Updates the two definitions of `add_library` that was missing the static
qualifier to ensure they generate the code we want.
2025-01-07 16:01:34 -08:00
Ryan Houdek a668492fb7 CPUID: Remove duplicated ARM Neoverse-N2
This was declared twice in the list.
2025-01-07 15:32:38 -08:00
Ryan Houdek da069571f3 Docs: Update for release FEX-2501 2025-01-07 13:07:46 -08:00
Ryan Houdek d2bac45b49 Merge pull request #4256 from bylaws/crtd
Windows: Only deinit the thread CRT when destroying the current thread
2025-01-06 21:46:32 -08:00
LC 8913c59acc Merge pull request #4250 from Sonicadvance1/staticanalysis
Just a few things picked up from static analysis
2025-01-06 19:11:38 -05:00
LC c3261b4aeb Merge pull request #4249 from Sonicadvance1/log_bad_fork_flags
LinuxSyscalls: Log unhandled clone3 fork flags
2025-01-06 19:11:02 -05:00
LC c7fb95aec5 Merge pull request #4248 from Sonicadvance1/fix_cefsimple
LinuxSyscalls: Ensure CSIGNAL is merged back in to flags for clone2
2025-01-06 19:10:18 -05:00
Billy Laws c00cef6dc1 WOW64: Fix warning 2025-01-06 19:07:07 +00:00
Billy Laws 429ff94dc5 Windows: Only deinit the thread CRT when destroying the current thread
The thread termination callback can be called for other threads in the
process, not just the current one, in which case we cannot call DeinitCRT.
Deinitializing the CRT of another thread would be awkward so just skip that
and accept the small leak for now.
2025-01-06 19:07:07 +00:00
LC a6c67ca749 Merge pull request #4251 from Sonicadvance1/ir_numelements_to_elementsize
IR: Change convention from number of elements to elementsize
2025-01-04 18:39:24 -05:00
LC f51812a670 Merge pull request #4253 from Sonicadvance1/minor_f80_opt
x87StackOptimizationPass: Minor opt to f80 fchs and fabs
2025-01-04 06:42:29 -05:00
Ryan Houdek 686294f1c4 InstcountCI: Update 2025-01-03 13:49:40 -08:00
Ryan Houdek a47ed105e7 x87StackOptimizationPass: Minor opt to f80 fchs and fabs
It's faster to load the f80 sign mask from our named vector constants
than synthesizing the values. Changes a 4 instruction sequence to
synthesize to be 1 load.
2025-01-03 13:47:10 -08:00
Ryan Houdek b2d579a268 OpcodeDispatcher: Assert on invalid size to LoadRegCachePair
Coverity scan
2025-01-03 11:06:30 -08:00
Ryan Houdek eb1050092f OpcodeDispatcher: Assert on invalid size to SelectPairAddressMode
Coverity scan
2025-01-03 11:05:38 -08:00
Ryan Houdek b3794f5541 OpcodeDispatcher: FEX_UNREACHABLE in programming error case
Coverity scan
2025-01-03 11:05:38 -08:00
Ryan Houdek 1ecfa3253d IR: Change convention from number of elements to elementsize
The IR stores elementsize, where the json was wanting number of
elements. While the IR Emitter function declaration always wanted
element size. This was causing us to do a little dance from ElementSize
-> Number of elements -> ElementSize. Just pass the ElementSize directly
instead of this bogus little dance.
2025-01-03 11:01:03 -08:00
Ryan Houdek 5daf007b6a OpcodeDispatcher: FEX_UNREACHABLE in programming error case
Coverity scan
2025-01-03 10:34:22 -08:00
Ryan Houdek 8efa5febd0 LinuxSyscalls: Log unhandled clone3 fork flags
Make sure to pass the clone3 arguments all the way to the fork handler
so it can check the flags. Currently nothing I know of uses fork plus
the new clone3 flags, but it would be hard to see without any logging.
2025-01-03 09:03:47 -08:00
Ryan Houdek 5fee8028cd LinuxSyscalls: Ensure CSIGNAL is merged back in to flags for clone2
This fixes #4247
2025-01-03 08:34:44 -08:00
LC 6bc7a83c64 Merge pull request #4245 from Sonicadvance1/update_kernel_minspec
FEXLoader: Increase minimum kernel requirement from 5.0 to 5.15
2025-01-02 14:48:26 -05:00
LC e55b5d0d11 Merge pull request #4246 from Sonicadvance1/fix_typo
Linux: Fixes typo in removing RESOLVE_IN_ROOT flag
2025-01-02 14:46:53 -05:00
Ryan Houdek 19de7f2785 Linux: Fixes typo in removing RESOLVE_IN_ROOT flag 2025-01-02 10:18:07 -08:00
LC e32c5384ab Merge pull request #4243 from Sonicadvance1/fix_4155
FEXLoader: Enable early logs output to stderr
2025-01-01 14:23:51 -05:00
LC b391fe6b92 Merge pull request #4244 from Sonicadvance1/fix_4150
unittests/ASM: Fix incorrect instruction form test
2025-01-01 14:23:04 -05:00
Ryan Houdek 4cfb81156f FEXLoader: Increase minimum kernel requirement from 5.0 to 5.15
Brought up in #4225 where it had issues with Openat2 which was added in
5.8.

The main driving force around minimum kernel version requirement is that
the lowest kernel version in our CI is 5.15. A benefit to this choice is
that this is an LTS release, which is also what Ubuntu 22.04 is
shipping.

Once the single CI machine is fixed to ship something newer then the
next logical choice would be kernel 6.1 which is also LTS, but until
then just lift it to 5.15. This version was released in October 2021,
and is supported by the kernel developers until 2026. Our previous
minimum of 5.0 was released in March 2019, so a two year leap here.

This removes the openat2 workaround that was necessary to pass our CI
since it is no longer necessary.
2025-01-01 11:22:54 -08:00
Ryan Houdek 6121708e55 unittests/ASM: Fix incorrect instruction form test
This test was generating the wrong form of instruction. There's no way
to choose this form with nasm deliberately, so manually encode it.

Fixes #4150
2025-01-01 10:12:51 -08:00
Ryan Houdek 6ab214adea FEXLoader: Enable early logs output to stderr
Some early FEXServer startup log failures weren't getting printed
correctly. They were going through the LogManager but before FEXServer
setup, or even stderr/stdout logman setup. So they were just getting
written to -1 and failing.

Fixes #4155
2025-01-01 10:00:22 -08:00
LC 90b1ac4162 Merge pull request #4241 from Sonicadvance1/fix_h0f3a_rex_decode
OpcodeDispatcher: Fixes FEX's H0F3A table handling of REX.W
2025-01-01 11:55:08 -05:00
LC 3abe6c14a1 Merge pull request #4240 from Sonicadvance1/3dnow_modrm_sib_test
unittests: Adds a 3DNow! ModRM SIB encoding test
2025-01-01 11:53:11 -05:00
LC fc1b500eff Merge pull request #4242 from Sonicadvance1/missing_tests
unittests/ASM: Adds missing MMX PADDQ test
2025-01-01 11:52:15 -05:00
Ryan Houdek 5d47b9195b unittests/ASM: Adds missing MMX PADDQ test 2025-01-01 08:22:38 -08:00
Ryan Houdek a8272b74f6 unittests/ASM: Ensure REX.W prefixed instructions from H0F3A are tested
We just want to ensure these instructions are decoded, the regular tests
are ensuring that the behaviour is correct.
2025-01-01 08:22:19 -08:00
Ryan Houdek 12dc16780f OpcodeDispatcher: Fixes FEX's H0F3A table handling of REX.W
Most of this table ignores REX.W, but two encodings change behaviour
based on REX.W. These two encodings are PEXTRD/PEXTRQ and PINSRD/PINSRQ.

For every other instruction encoding, they will ignore REX.W, but FEX
was requiring that they didn't have REX.W encoding. I had special cased
this in the past by adding PALIGNR, but that didn't handle any of the
other instructions.

We can't just handle REX.W in the OpcodeDispatcher and remove the two
special cased instructions because these vector operations also interact
with instruction prefix 0x66 which changes the operating size to 16bit
with regular instructions.

So instead just generate all listings of instructions with REX.W being
zero and one and install handlers in all cases.
2025-01-01 08:22:19 -08:00
Ryan Houdek b8af569841 unittests: Adds a 3DNow! ModRM SIB encoding test
This codepath was unttested in our CI.
2025-01-01 08:21:56 -08:00
LC 8bee101795 Merge pull request #4232 from Sonicadvance1/disable_gvisor_tests
unittests/gvisor: Disable memfd tests
2025-01-01 08:38:36 -05:00
Ryan Houdek 2d66bc258a Merge pull request #4225 from asahilina/merged-rootfs
Support a merged RootFS (and a bunch of related fixes)
2024-12-31 17:29:06 -08:00
Ryan Houdek d2f86e49f7 Merge pull request #4237 from bylaws/fpfix
Fix float->int conversion overflow behaviour
2024-12-31 16:00:20 -08:00
Ryan Houdek d66cd16cfb Merge pull request #4230 from asahilina/thunks-build-sysroot
Library Forwarding: Allow reading standard library headers from a development x86 rootfs
2024-12-30 18:00:34 -08:00
Ryan Houdek 04e785e434 Merge pull request #4231 from Sonicadvance1/minor_div_opt
OpcodeDispatcher: Minor division improvement
2024-12-30 17:32:53 -08:00
Ryan Houdek 15a1a0f7d9 Merge pull request #4239 from bylaws/3dn
Frontend: Fix ModRM handling with 3DNow!
2024-12-30 17:31:58 -08:00
Billy Laws 0a58ce6134 Frontend: Fix ModRM handling with 3DNow! 2024-12-30 18:35:39 +00:00
Billy Laws 8f5607f0e8 Update InstCountCI 2024-12-30 01:07:36 +00:00
Billy Laws a21789d3d8 ASM_Tests: Test F2I conversion overflow behaviour 2024-12-30 00:47:00 +00:00
Billy Laws efd6e95059 OpcodeDispatcher: Match x86 overflow behaviour for F2I conversions
ARM behaviour here is to saturate on overflow or NaN inputs, whereas
X86 returns a sentinel value of 2^(bitsize-1), explicitly emulate this.
2024-12-30 00:42:55 +00:00
Billy Laws 9bdb1f4306 OpcodeDispatcher: Make narrowing implicit for F64->I32 conversions
This is always used, removing it avoids needing to handle unused codepaths.
2024-12-30 00:36:17 +00:00
Billy Laws ae4b7135d5 OpcodeDispatcher: Share AVX F2I/I2F code for 256-bit SVE 2024-12-30 00:29:31 +00:00
Tony Wasserka d503366816 Library Forwarding: Allow reading standard library headers from a development x86 rootfs 2024-12-24 19:41:29 +09:00
Ryan Houdek 0fe2827fcc unittests/gvisor: Disable memfd tests
This tests some bugged or changed behaviour. So we need to disable these
since our CI crosses kernel versions that hit both behaviour paths.
2024-12-22 03:11:08 -08:00
LC cd6722f77b Merge pull request #4229 from Sonicadvance1/more_lrcpc2_tests
InstCountCI: Adds more LRCPC2 tests that are missed
2024-12-20 22:57:06 -05:00
Ryan Houdek ffb745b662 InstCountCI: Update for divison improvements 2024-12-20 13:22:42 -08:00
Ryan Houdek bb10f25808 OpcodeDispatcher: Minor division improvement
No need to extract the subregisters out before operating on them since
the long division and long remainder IR operations correctly zero/sign
extend the incoming sources as necessary. Saves a couple of
instructions.
2024-12-20 13:20:52 -08:00
Ryan Houdek aa1076d12b InstCountCI: Adds more LRCPC2 tests that are missed
We weren't testing 64-bit variants, and we also weren't testing 8-bit
and 16-bit loadstores. Add some more to ensure we are hitting these.
2024-12-20 12:12:24 -08:00
Ryan Houdek 1e827ec7a6 Merge pull request #4227 from Sonicadvance1/fix_atomic_loadstore
ArchHelpers/Arm64: Fixes LDAPUR and STLUR backpatching
2024-12-20 11:46:13 -08:00
Asahi Lina 3fe2650787 FileManagement: Gate new openat2() codepaths on recent enough kernel 2024-12-21 00:52:12 +09:00
Asahi Lina 3e99e814bc FileManagement: Use openat2() with RESOLVE_IN_ROOT for RootFS open ops
This avoids having to do the symlink chasing in GetEmulatedFDPath, since
the kernel does it for us. On top of that, with a merged RootFS
setup, this will correctly handle symlinks from user directories into
the RootFS, fixing wine on Fedora.
2024-12-21 00:52:11 +09:00
Ryan Houdek 2019f8138e ArchHelpers/Arm64: Fixes LDAPUR and STLUR backpatching
The immediate offset masking was at the completely wrong offset when I
wrote these handlers. No idea how I managed to mess those up so badly.

Should fix at least some of the issues with #4216
2024-12-19 17:29:45 -08:00
LC e44d1f136b Merge pull request #4226 from alyssarosenzweig/instc/factorio
InstructionCountCI: add some hot blocks from Factorio
2024-12-19 15:52:59 -05:00
Alyssa Rosenzweig 09872402df InstructionCountCI: add some hot blocks from Factorio
Factorio hammers its drawSprite() function and ends up cpu bound under FEX.
Unfortunately, its hot blocks seem to be translated pretty optimally :-/

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-12-19 15:15:38 -05:00
Asahi Lina b078a41a02 FileManagement: Fix return val of readlink*
The wrappers handle errno, we just need to return -1 on errors.
2024-12-20 03:01:54 +09:00
Asahi Lina 3a5eeb5700 Syscalls: Fix multiple shebang handling issues
- Parse the shebang line properly (use FHU::ParseArgumentsFromString
  which is the same code the loader uses)
- Make native-interpreter shebang files work by deferring to the kernel
  in that case (previously, they'd get executed through the loader and
  it would choke on the architecture of the interpreter)
- Do not use the RootFS-prepended path when executing shebang files. The
  loader will prepend that anyway when looking it up, but it needs the
  bare guest path so it can pass it as an argument to the interpreter,
  which (since it's emulated) will do the lookup through the RootFS.
2024-12-20 03:01:54 +09:00
Asahi Lina 9433ae3405 Syscalls: Handle execve of native binaries with merged RootFS
With a merged RootFS, all binaries are executed through the RootFS. When
executing a binary that is actually a native binary, we want to do so
outside the RootFS. Handle this by stripping the RootFS prefix in that
case.
2024-12-20 01:58:12 +09:00
Asahi Lina 4658b24f9a FileManagement: Handle RootFS symlinks into RootFS properly
If a RootFS symlink links to an absolute path within the RootFS, we need
to strip the RootFS prefix. This would not normally happen with a plain
RootFS, but it can happen if /proc is mounted within the RootFS.
2024-12-20 00:41:07 +09:00
Asahi Lina 4e7d0e6be0 FileManagement: Fix path resolution for symlinks to the root
If there's a symlink to / within the RootFS, don't attempt to follow it,
since that will end up trying to look up the empty string within the
RootFS (which is not legal). Just return the symlink.
2024-12-20 00:41:07 +09:00
Asahi Lina 4ddd98708f FileManagement: Handle readlink /proc/self/fd/* properly
If the guest reads a RootFS path from /proc/self/fd/*, we should return
it with the RootFS prefix stripped.
2024-12-20 00:41:07 +09:00
Asahi Lina c161fd218c FileManagement: Simplify emulated file lookup
To locate whether a path is in the emulated list, EmulatedFDManager::OpenAt()
attemps to resolve the path. realpath() ends up calling readlinkat() on
every path component, which is a lot of syscalls for every open()
variant syscall. It also makes interaction with the rootfs complex and
error-prone.

There's a much easier way to do this: We just open the file without
emulation and check its real path via get_fdpath(). This is just one
readlink() syscall per open, instead of one per path component. If the
file turns out to be emulated (uncommon case), we swap out the fds.

This also decouples EmulatedFDManager from guest path resolution
entirely, so it will never fall out of sync with the RootFS logic.
2024-12-20 00:41:07 +09:00
LC 7e257cc268 Merge pull request #4222 from bylaws/fmtt
External: Update bundled libfmt
2024-12-18 19:54:55 -05:00
Ryan Houdek d8ef70280c Merge pull request #4221 from Sonicadvance1/threadmanager_footexplosions
ThreadManager: Add some sanity asserts
2024-12-18 11:30:19 -08:00
Billy Laws ec003281be External: Update bundled libfmt 2024-12-18 15:25:45 +00:00
Ryan Houdek e58f67b76c ThreadManager: Add some sanity asserts
These couple of functions have some footguns that I'm encountering while
rewriting gdbserver. Ensure that assertion builds capture the problems
2024-12-17 15:06:53 -08:00
LC 57178abcd2 Merge pull request #4220 from Sonicadvance1/expose_faultsafe
Linux/FaultSafeUserMemAccess: Break out fault safe handler
2024-12-16 17:02:06 -05:00
Ryan Houdek 73ca4f8314 Linux/FaultSafeUserMemAccess: Break out fault safe handler
This is going to get used by gdbserver soon for ensuring memory accesses
are fault safe, because it tries to read outside of correct memory
bounds at times.
2024-12-16 11:06:15 -08:00
LC 527752c25b Merge pull request #4218 from Sonicadvance1/fix_file_loading
Utils/FileLoading: Fix LoadFileImpl
2024-12-13 22:57:35 -05:00
Ryan Houdek 38fa866c91 Utils/FileLoading: Fix LoadFileImpl
It is not an error that pread returns /less/ than what was requested. In
fact it's very common for the Linux kernel to return less than the data
requested from procfs.

procfs keeps coming back to bite this function, previously it was fstat
returning size of 0 which it hit. Now it only feeds data as much as it
wants per loop. In particular /proc/self/maps would only read ~3k bytes
on my system, but not be complete.

To fully fix the issue, always make sure to keep reading until there is
either an error OR zero is reached!
2024-12-13 19:42:00 -08:00
Ryan Houdek c902b8807a Merge pull request #4215 from alyssarosenzweig/fix/constprop-zext
ConstProp: fix 32-bit masking behaviour
2024-12-13 17:33:30 -08:00
Alyssa Rosenzweig 4934c1fd94 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-12-13 10:44:56 -05:00
Billy Laws 766fbe3db3 unittests: Add a test for constprop size bugs
fails on main, fixed by this PR.
2024-12-13 10:44:56 -05:00
Alyssa Rosenzweig 29405f2690 ConstProp: fix 32-bit masking behaviour
if we want to replace a node with one of its sources, we need to zero extend if
the source is 64-bit and the destination is 32-bit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-12-13 10:44:56 -05:00
Alyssa Rosenzweig 51f505acca ConstProp: drop some unused headers
ycm complained.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-12-13 10:44:56 -05:00
Alyssa Rosenzweig 77415538f7 OpcodeDispatcher: use 64-bit XOR for AF calc
we don't need masking and the masking gets in the way of constprop.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-12-13 10:44:56 -05:00
Alyssa Rosenzweig 9fb69ed206 Merge pull request #4209 from Sonicadvance1/tso_support_instcountci
InstCountCI: Implement support for TSO and LRCPC and add hot block that could be optimized
2024-12-13 09:33:45 -05:00
LC 735a4f90db Merge pull request #4212 from Sonicadvance1/fix_encoding
GdbServer: Fixes encoding of hex
2024-12-12 22:49:43 -05:00
Ryan Houdek 7ef8dc13ba GdbServer: Fixes encoding of hex
Just a typo accidentally prefixing 0x on the hex when it shouldn't.
2024-12-12 16:15:41 -08:00
Ryan Houdek 656477ec63 Merge pull request #4165 from bylaws/denuvo
Support inline self modifying code
2024-12-12 13:38:31 -08:00
Billy Laws d080180e85 ARM64EC: Process pending cross-process work on syscalls and exceptions
This is used to notify the JIT of e.g. memory writes by a debugger.
2024-12-12 21:28:37 +00:00
Billy Laws af1d2d6005 ARM64EC: Implement inline SMC support using context reconstruction
When an SMC trap happens: reconstruct the context before the SMC write
then compile the write as a single instruction block to reduce it to
regular SMC. SMC where the writing instruction is the instruction being
patched will hit the signal handler at most twice: the 1st will trigger
the write to be compiled as a single instuction block, the 2nd will
detect inline SMC of a single instruction block and then just take the
usual invalidate+reprotect+continue step, avoiding a potential infinite
loop of recompilation.
2024-12-12 21:28:37 +00:00
Billy Laws 90c1282f3a Dispatcher: Support forcing a temp single instr block on ARM64EC JIT entry 2024-12-12 21:28:37 +00:00
Billy Laws d5d7eec8b0 FEXCore: Expose an API to check if the current block represents a single
guest instruction

Single instruction blocks need to be treated specially when inline SMC
is detected, the frontend only needs to reprotect RWX and invalidate
caches then continue execution as side effects from the SMC shouldn't be
seen until the instruction executes.
2024-12-12 21:28:37 +00:00
Billy Laws 5337b9537d FEXCore: Expose an API to query intersection with the current block
Frontends need to detect this in order to handle SMC within the current
block (inline SMC) differently to regular SMC which can just reprotect
and continue.
2024-12-12 21:28:37 +00:00
Billy Laws e72c016230 Core: Split blocks on invalid instructions 2024-12-12 21:28:37 +00:00
Ryan Houdek 072cf4c5bd Merge pull request #4205 from Sonicadvance1/gdbserver_support_32bit
GdbServer: Support 32-bit context definitions
2024-12-12 12:55:24 -08:00
Ryan Houdek 27ededf47f Merge pull request #4206 from bylaws/smcim
Windows: Track RWX regions in mapped images
2024-12-12 12:53:13 -08:00
Ryan Houdek 82d7f9fdd7 GdbServer: Support 32-bit context definitions
Requires restructuring a couple of things, but nothing too crazy here.
2024-12-12 12:35:58 -08:00
Ryan Houdek d85153d6b3 GdbServer: Save off some signal information when it occurs
Enough for some state reconstruction that is missing
2024-12-12 12:14:55 -08:00
Ryan Houdek 6b698e6cd1 SignalDelegator: Make SpillSRA public
GdbServer wants to use it
2024-12-12 12:14:54 -08:00
Ryan Houdek 9475f79ec6 GdbServer: Save off SignalDelegator 2024-12-12 12:14:54 -08:00
Ryan Houdek f906c6a0f4 Merge pull request #4211 from asahilina/pthread-attr-memleak
Threads: Fix memory leak in joinable()
2024-12-12 12:13:17 -08:00
Ryan Houdek e88c92de57 Merge pull request #4161 from bylaws/tf
FEXCore: Emulate EFLAGS.TF
2024-12-12 11:51:53 -08:00
Asahi Lina 48ed906a7b Threads: Fix memory leak in joinable() 2024-12-13 04:47:41 +09:00
LC b03b02d2f2 Merge pull request #4210 from Sonicadvance1/add_missing_comment
IR/Passes: Adds missing comment that clang-format keeps complaining about locally
2024-12-11 18:51:57 -05:00
Ryan Houdek d00d476a0a IR/Passes: Adds missing comment that clang-format keeps complaining about locally
NFC
2024-12-11 15:03:28 -08:00
Ryan Houdek ac1e32994a InstCountCI: Adds hot block that doesn't generate optimal code 2024-12-11 15:03:02 -08:00
Ryan Houdek 800d447f3d InstCountCI: Add support for TSO and LRCPC1/2 2024-12-11 14:55:19 -08:00
LC 8111b7cc7f Merge pull request #4194 from Sonicadvance1/fcw_pc_instructions
FEXCore: Override x87 precision control when necessary
2024-12-10 17:24:55 -05:00
LC a86c922073 Merge pull request #4203 from Sonicadvance1/const_reconstruct
Context: Constify GPRs passed to ReconstructCompactedEFLAGS
2024-12-10 12:46:05 -05:00
LC 46fb8583bb Merge pull request #4204 from Sonicadvance1/gdbserver_vkill
GdbServer: Implement support for `$vKill`
2024-12-10 12:45:06 -05:00
Billy Laws 3487d120ec Windows: Treat PAGE_EXECUTE_WRITECOPY memory as RWX 2024-12-10 15:26:23 +00:00
Billy Laws 07394d6a6e Windows: Track RWX regions in mapped images
As section permissions are set on the unix side we don't get a
protection callback for them, workaround this by iterating over
the sections of all executables after mapping and tracking the RWX
ones.
2024-12-10 15:24:54 +00:00
Billy Laws 8d3204171c instcountci: update 2024-12-10 15:24:03 +00:00
Billy Laws 7641f722e9 unittests: Test TF 2024-12-10 15:20:47 +00:00
Billy Laws b51fa497c5 OpcodeDispatcher: Mask TF for pop ss instructions 2024-12-10 15:20:47 +00:00
Billy Laws 34722bed3d SignalDelegator: Clear TF when running signal handlers 2024-12-10 15:20:47 +00:00
Billy Laws 981c3009ee FEXCore: Emulate EFLAGS.TF
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
2024-12-10 15:20:47 +00:00
Ryan Houdek 38cf357d85 GdbServer: Implement support for $vKill
This is the command used when the `k` argument is passed to gdb. There
is nothing to do once this is received other than "kill" as quickly as
possible. The absolute way to ensure this is using SIGKILL.

No way to do a `r` command after `k` yet, but might be possible.
2024-12-09 15:13:52 -08:00
Ryan Houdek 2533ed4a63 Context: Constify GPRs passed to ReconstructCompactedEFLAGS
This only reads the GPRs passed in, doesn't modify it.
2024-12-09 15:08:59 -08:00
Ryan Houdek 5a4691fdfc Merge pull request #4201 from Sonicadvance1/remove_lock
FEXCore: Don't `WaitForEmptyJobQueue` if CodeObjectCacheService isn't used
2024-12-09 10:35:20 -08:00
Billy Laws 6c035a0d61 Dispatcher: Split out some common code into lambdas 2024-12-09 14:15:28 +00:00
Billy Laws a234aa300d Dispatcher: Skip extra L1 lookup after CompileBlock 2024-12-09 12:30:00 +00:00
Billy Laws f6abbedbd1 ARM64EC: Fix typo so TF is unset handling exceptions 2024-12-09 12:30:00 +00:00
Billy Laws b6fe4cd6dd Windows: Skip state reconstruction on exceptions in dispatcher
The dispatcher always spills register state before issuing faulting instructions
2024-12-09 12:30:00 +00:00
LC bdae4f6915 Merge pull request #4200 from Sonicadvance1/fix_exit
LinuxSyscalls: Fixes exit syscall
2024-12-08 19:31:16 -05:00
LC f8b6edfb2b Merge pull request #4199 from Sonicadvance1/remove_arch
docs: Remove Arch from the release process.
2024-12-08 19:29:42 -05:00
Ryan Houdek 0a1ecdf6ae FEXCore: Don't WaitForEmptyJobQueue if CodeObjectCacheService isn't used
Seems the unused mutex locking is able to cause some hangs according to #4198
Hard to tell why, but might as well as get rid of that potential
pitfall.
2024-12-08 08:01:28 -08:00
Ryan Houdek beec203f56 LinuxSyscalls: Fixes exit syscall
if an application is using `exit` then it is usually a faulting
condition rather than cleanly exiting. When cleanly exiting
applications will typically use `exit_group` instead.

`exit` is useful to quickly cause a single thread to exit in a
multi-threaded environment as well, where `exit_group` will take down
the entire process group.

FEX had implemented this in a way that would do a double Stop signal,
cascading to a crash. When tied in to a crash handler, this could get
caught in a weird way.

This /should/ fix #4198, but I can't confirm locally. It looks like in
that issue that the steam install is slightly buggered (as evident by
missing srt-logger and steam-runtime-identify-library-abi).

This is a bug regardless so fix it and create a unittest. If it doesn't
fix the user's bug, then we have another workaround that will definitely
solve it.
2024-12-08 05:14:19 -08:00
Ryan Houdek d323032ec9 docs: Remove Arch from the release process.
On December 6th 2024, the fex-emu packages got a deletion request:

> MarsSeed [1] filed a deletion request for fex-emu [2]:
>
> ARM-only package.
> This should be submitted to ArchLinuxARM.org [a], not to AUR - see
> quote from ArchWiki [b]:
>
>     "Packages that do not support the x86_64 architecture
>     are not allowed in the AUR."
>
> [a]:
> https://archlinuxarm.org/forum/viewforum.php?f=4
> [b]:
> https://wiki.archlinux.org/title/AUR_submission_guidelines#Rules_of_submission
>
> [1] https://aur.archlinux.org/account/MarsSeed/
> [2] https://aur.archlinux.org/pkgbase/fex-emu/

This is due to a rule clarification that occured in Arch's forum on November 25th: https://lists.archlinux.org/archives/list/aur-general@lists.archlinux.org/thread/IRZ2LWYX3ECPJQZJXMLAP6JIKL6HLHPZ/#GMYC74CRSFH7GGNENEUOODZUPWHOMX7A

On December 3rd the package submission guidelines on their wiki was
updated to mandate x86-64 support:
https://wiki.archlinux.org/index.php?title=AUR_submission_guidelines&diff=prev&oldid=822050

As of today, December 7th, 2024 the packages have been removed from AUR
due to only supporting aarch64.

> Muflone [1] deleted fex-emu [2].
>
> You will no longer receive notifications about this package.
>
> [1] https://aur.archlinux.org/account/Muflone/
> [2] https://aur.archlinux.org/pkgbase/fex-emu/

ArchLinux is no longer a supported distro for FEX, remove it from the release processes documentation.
2024-12-07 16:15:22 -08:00
Ryan Houdek 7472b21f33 Merge pull request #4197 from Sonicadvance1/revert_4118
Revert #4118
2024-12-07 11:39:39 -08:00
Ryan Houdek 1058575d3a InstcountCI: Update pause instruction 2024-12-06 17:04:28 -08:00
Ryan Houdek e9867ca35a Revert "FEXCore: Change yield implementation to use wfe"
This reverts commit e53f3969e9.
2024-12-06 17:02:27 -08:00
Ryan Houdek 84277319fa Merge pull request #4166 from pmatos/HostFeaturesInPass
Generate SVE for 80bit load/stores when possible
2024-12-06 02:01:07 -08:00
Paulo Matos 8f8aa55c7f instcountci: Cache predicate register generation from pattern 2024-12-06 10:15:38 +01:00
Paulo Matos 72a4063651 Cache predicate register generation from pattern 2024-12-06 10:15:38 +01:00
Paulo Matos 0b1229da55 instcountci: Generate SVE for 80bit load/stores when possible 2024-12-06 10:15:38 +01:00
Paulo Matos 1d3ce30e50 Generate SVE for 80bit load/stores when possible
Fixes #4166.
2024-12-06 10:15:29 +01:00
LC 71187d3ad7 Merge pull request #4195 from Sonicadvance1/fix_clone3
LinuxEmulation: Don't use clone3 for fork
2024-12-06 00:10:07 -05:00
Ryan Houdek dd8a3a9aea LinuxEmulation: Don't use clone3 for fork
clone3 was added in Linux 5.3 but our minimum spec is 5.0. Additionally
the Raspberry Pi 5 kernel seems to complain about clone3 for some
reason?

Just use clone instead of clone3
2024-12-05 15:14:37 -08:00
Tony Wasserka 7b2fc37651 Merge pull request #4193 from WhatAmISupposedToPutHere/main
Thunks/gen: Add support for compiling against clang 19
2024-12-05 15:35:06 -05:00
Sasha Finkelstein 426569d74d Thunks/gen: Add support for compiling against clang 19 2024-12-05 21:16:41 +01:00
Ryan Houdek e877d5b82c unittests: Disable failing x87 tests on simulator 2024-12-05 00:03:33 -08:00
Ryan Houdek 572e0d04d5 unittests/X87: Adds precision and rounding mode tests
Tests all the instructions that are affected by FCW PC (or not!)
Only missing tests are fsincos (More easily tested with just fsin and
fcos), and fpatan
2024-12-04 23:54:36 -08:00
Ryan Houdek e3d7161ac5 FEXCore: Override x87 precision control when necessary
According to the documentation for x87 FCW precision control, this only
affects fadd*, fsub*, fmul*, fdiv*, and fsqrt. FEX was incorrectly
reducing precision for all x87 operations.

Precision is ignored for the following x87 ALU operations:
- fabs
- fscale
- fprem{1,}
- fcos
- fsin
- ftan
- fyl2x
- fyl2xp1
- fpatan
- fsincos
- Plus any operations just doing data movement and conversions

Next commit adds unittests to ensure this is correct for each
instruction.
2024-12-04 23:47:19 -08:00
Ryan Houdek 20caf69951 Docs: Update for release FEX-2412 2024-12-03 11:02:44 -08:00
Paulo Matos fcbf0de05a Enable RA of SVE Predicate Registers 2024-12-02 18:35:31 +01:00
Ryan Houdek 731e4d6271 Merge pull request #4168 from pmatos/SVEStoreInstCountCI
instcountci: testing multiple 80bit ldst using SVE
2024-12-02 07:56:12 -08:00
Paulo Matos 48beb18f29 instcountci: testing multiple 80bit ldst using SVE
In preparation for #4166 which should improve on these results.
2024-12-02 11:37:11 +01:00
LC 41c8731443 Merge pull request #4188 from Sonicadvance1/fexcore_remove_unnecessary
FEXCore: Removes ExitHandler and RunUntilExit
2024-12-01 15:06:57 -05:00
Ryan Houdek efb276f489 FEXCore: Removes ExitHandler and RunUntilExit
Now that all the threading behaviour has been correctly separated/moved
to the frontend, these functions serve no purpose.

- Instead of using RunUntilExit, all threads can use `ExecuteThread`
  directly, since there's nothing special about the primary thread now.
  - This also removes the public function definition of `ExecutionThread` since that was only used for threading logic.
- Instead of using an exit handler, just do the same cleanup after
  `ExecuteThread` has returned.
  - Just make gdbserver is cleaned up early if it exists since it may
    want to send some things to the connected gdb instance before
    threads are exited.
2024-12-01 10:45:38 -08:00
LC 9febddefa3 Merge pull request #4190 from Sonicadvance1/remove_old_gprsize
FEXCore: Removes GetGPRSize and convert all uses to GetGPROpSize
2024-12-01 13:13:21 -05:00
Ryan Houdek ce9a860335 OpcodeDispatcher: Also remove unused CacheIndexToSize 2024-12-01 05:38:11 -08:00
Ryan Houdek 0123946ed1 FEXCore: Removes GetGPRSize and convert all uses to GetGPROpSize
Only a few remaining uses left, easy enough to convert. This finally
switches the final few uses over.

NFC
2024-12-01 05:38:09 -08:00
LC c7098d0da1 Merge pull request #4189 from Sonicadvance1/fexcore_remove_definition
FEXCore: Removes stale function definition
2024-12-01 08:29:11 -05:00
LC 65a162bdf9 Merge pull request #4187 from Sonicadvance1/fexcore_remove_coreshuttingdown
FEXCore: Removes CoreShuttingDown from ContextImpl
2024-12-01 08:28:32 -05:00
Ryan Houdek 2de485d02a FEXCore: Removes stale function definition
`CopyMemoryMapping` was removed a long time ago, the definition happened
to remain. Remove the definition.
2024-11-30 23:20:07 -08:00
Ryan Houdek bf64facaf6 FEXCore: Removes CoreShuttingDown from ContextImpl
This is unused now.

NFC
2024-11-30 22:10:37 -08:00
LC 2e7fc60dbf Merge pull request #4169 from Sonicadvance1/x87_loadstore_tests
unittests/ASM: Fixes x87 80-bit loads on the edge of page boundaries.
2024-11-30 10:43:10 -05:00
LC baddfe00b1 Merge pull request #4186 from Sonicadvance1/remove_remaining_runningevents
FEXCore: Removes remaining RunningEvents from InternalThreadState
2024-11-29 20:00:00 -05:00
Ryan Houdek e7e59204d3 FEXCore: Removes remaining RunningEvents from InternalThreadState
These are all frontend constructs with mostly deprecated constraints.
WaitingToStart isn't used anymore, Running is effectively always true
(and behaviour has changed that if a thread is alive, it's running).

The only one that remains is `ThreadSleeping` which is only handled in
the frontend, and there was some conflation between ThreadSleeping and
Running which was hard to gauge. So delete `Running` and
`WaitingToStart`, but move `ThreadSleeping` to the frontend.
2024-11-29 14:08:55 -08:00
Ryan Houdek 95d5b14f99 Merge pull request #4183 from Sonicadvance1/move_fexcore_executionthread
FEXCore: Moves InternalThreadState ExecutionThread to the frontend
2024-11-29 14:06:42 -08:00
Ryan Houdek 802eaee9c8 FEXCore: Moves InternalThreadState ExecutionThread to the frontend
Once again this is another frontend construct, so move it to
ThreadStateObject
2024-11-29 13:33:56 -08:00
LC e89f48f237 Merge pull request #4182 from Sonicadvance1/move_start_paused
FEXCore: Move InternalThreadState StartRunning to frontend
2024-11-29 16:31:37 -05:00
Ryan Houdek e771e25632 LinuxSyscalls/Thread: Build child thread arguments on parent stack
Now that most of the thread tracking is in the frontend, change this
over to building the thread execution handler on the parent thread.

Removes a memory allocation/free pair, and removes the copy of each
variable in the child thread.
2024-11-29 09:38:36 -08:00
Ryan Houdek 25c202575e FEXCore: Move InternalThreadState StartRunning to frontend
We were using this variable for two things, letting the frontend signal
to the backend that it wants to start executing once the thread is
created, and also for handling thread pausing. These two features are
conflated with one another and actually makes things more confusing.

- Move StartRunning/StartPaused to the frontend, because its a construct
  that only needs to exist in the frontend
- Adds a FEX::HLE::ThreadStateObject CV for handling pausing, which only
  needs to exist for gdbserver
2024-11-29 09:38:24 -08:00
Ryan Houdek f59fc0f747 Merge pull request #4181 from Sonicadvance1/remove_exitreason
FEXCore: Removes ExitReason from InternalThreadState
2024-11-29 09:36:33 -08:00
Ryan Houdek f7a076e00c FEXCore: Removes ExitReason from InternalThreadState
FEXCore hasn't been returning anything other than EXIT_SHUTDOWN for a
long time, so this ended up just moving data around for no reason.

This isn't going to be used for further GdbServer work anyway, so just
completely remove it.
2024-11-29 09:25:44 -08:00
Ryan Houdek 56fadecdaf Merge pull request #4179 from Sonicadvance1/move_thread_waiting_start
FEXCore: Moves ThreadWaiting to the frontend
2024-11-29 09:23:26 -08:00
Ryan Houdek 9f681f9e41 FEXCore: Moves ThreadWaiting to the frontend
Only in one location does the frontend actually care about this, the
backend doesn't care at all.
2024-11-29 09:10:36 -08:00
Ryan Houdek 1b11f2f184 Merge pull request #4185 from asahilina/fix-autoshutdown-regression
FEXServer: Fix auto-shutdown regression
2024-11-29 09:03:53 -08:00
LC b2e61c37be Merge pull request #4170 from Sonicadvance1/gdbserver_work
GdbServer: Minor work
2024-11-29 08:24:43 -05:00
Asahi Lina 7c0cf51f09 FEXServer: Fix auto-shutdown regression
Fixes: #4184
2024-11-29 20:56:07 +09:00
Ryan Houdek 649a49488b Merge pull request #4180 from Sonicadvance1/fexcore_const_ptr_ctx
FEXCore: Constify CTX ptr in InternalThreadState
2024-11-28 17:02:16 -08:00
Ryan Houdek aa2180d494 Merge pull request #4178 from Sonicadvance1/move_statuscode_frontend
FEXCore: Moves StatusCode to the frontend
2024-11-28 16:51:33 -08:00
Ryan Houdek 969cae581c FEXCore: Constify CTX ptr in InternalThreadState
The CTX pointer in the InternalThreadState object will not and must not
change, since it is associated with that CTX object.

Contify it to codify it.
2024-11-28 15:57:36 -08:00
Ryan Houdek 1bf7e2544a FEXCore: Moves StatusCode to the frontend
This is a Linux construct, move it to the frontend.

This is going to need some changes in the future since exit_group and
exit syscalls are supposed to behave differently than how FEX implements
it. For now just move it to the frontend.
2024-11-28 15:55:46 -08:00
Ryan Houdek fad22144a2 Merge pull request #4177 from Sonicadvance1/move_deferred_signal_state
FEXCore: Moves DeferredSignalFrames to the frontend
2024-11-28 15:55:02 -08:00
Ryan Houdek b440e176fb Merge pull request #4176 from Sonicadvance1/move_signalreason
FEXCore: Moves SignalThread/SignalEvent to Frontend
2024-11-28 15:54:21 -08:00
Ryan Houdek 56c6b0d2cb Merge pull request #4175 from Sonicadvance1/gdbserver_remove_earlyexit
FEXCore: Removes EarlyExit running event
2024-11-28 15:53:17 -08:00
Ryan Houdek 0596a963e1 Merge pull request #4174 from Sonicadvance1/gdbserver_move_alloc_tls
FEXCore: Moves TLS initialization for Alloc::OSAllocator
2024-11-28 15:49:19 -08:00
Ryan Houdek 357cc04940 GdbServer: Splits Multi-letter v command handler
Just breaks out the two commands we support and leaves TODOs for
implementing the remaining commands.

NFC
2024-11-28 15:32:50 -08:00
Ryan Houdek 7c6e836865 GdbServer: Split out GDB context definition generation to its own function
GDB has two ways to read the registers. One way is reading the full
GDBContextDefinition, which matches the layout in `BuildTargetXML`.

The other way is to read the individual elements out of
GDBContextDefinition.

These two code paths were independently implemented. Instead generate in
one location and use in either location.

NFC
2024-11-28 15:32:50 -08:00
Ryan Houdek 54a7317312 GdbServer: Split out function searching for thread by TID
This currently happens in two locations, so split it out.

There's some behaviour here where if the TID isn't found, then it
returns the ParentThread of the process. This is working around a bug in
either FEX's gdbserver or binaryninja. Leave it currently before we
figure out what's wrong.

NFC
2024-11-28 15:32:50 -08:00
Ryan Houdek 1fb20710e6 GdbServer: Switch to thread specific stopping break logic
Previous `S AA` logic is legacy for non-multithreaded applications. This
newer command gives more information about what occured and in what
thread id.
2024-11-28 15:32:50 -08:00
Ryan Houdek 811ea093b5 GdbServer: Split out qXfer handlers
NFC, just making this easier to track for me.
2024-11-28 15:32:50 -08:00
Ryan Houdek 71fe9aee21 GdbServer: Fixes thread name setting
When parsing `comm`, by default it will have a newline which breaks gdb
in some cases. Strip out the whitespace to fix that issue.
2024-11-28 15:32:50 -08:00
Ryan Houdek 38c834e731 FEXCore: Removes global StartPaused check for gdb
This doesn't behave properly anymore now that thread management was
moved to the frontend.
2024-11-28 15:32:50 -08:00
Ryan Houdek 740ff60a71 GdbServer: Reorganize packet command handlers
Makes these consistent in the handling and documents the commands in a
way that is easier to parse while working on this.

NFC
2024-11-28 15:32:50 -08:00
Ryan Houdek ee69b9f650 GdbServer: Reconstruct XMM/YMM registers using FEXCore helpers
Previously this would have corrupted data in the upper 128-bits of the
YMM register.
2024-11-28 15:32:50 -08:00
Ryan Houdek a4565ce783 GdbServer: Pass through FCW
We have supported this for a while, just wasn't passed through gdbserver
since it usually doesn't matter.
2024-11-28 15:32:50 -08:00
Ryan Houdek 3131ee4de1 GdbServer: Moves information fetching to independent files
NFC
2024-11-28 15:32:50 -08:00
Ryan Houdek 3ecc66fbcf FEXCore: Moves TLS initialization for Alloc::OSAllocator
Alloc::OSAllocator uses a TLS variable of the thread object so it can
use a forkable mutex plus a deferring signal section. This was setup
when the FEXCore "ExecutionThread" function is called, which is a bit
awkward and is an artifact from when the thread creation was mixed
between the frontend and the backend.

Instead let the frontend inform the backend when to install the TLS
variable.

This is one step required to make GdbServer work correctly again since
the thread initialization and pausing is awkward today.
2024-11-28 02:46:06 -08:00
Ryan Houdek e322e84785 FEXCore: Moves DeferredSignalFrames to the frontend
Deferred signal frames are a frontend construct. Move it there.
2024-11-28 01:14:55 -08:00
Ryan Houdek f0fa7a5b6a FEXCore: Moves SignalThread/SignalEvent to Frontend
This is purely a Linux frontend construct now, move it.
2024-11-28 00:57:50 -08:00
Ryan Houdek 718221be71 FEXCore: Removes EarlyExit running event
This was working around an edge case in the GdbServer where a thread was
getting created while the process was shutting down. This edge case is
getting removed so get rid of it.
2024-11-28 00:46:42 -08:00
Ryan Houdek ee592ba03c Merge pull request #4173 from neobrain/fix_ctest_list
CMake: Generate test list even when testing is disabled
2024-11-27 13:52:32 -08:00
Tony Wasserka 56947f3a94 CMake: Generate test list even when testing is disabled
Previously, running ctest with BUILD_TESTS=OFF would discover and execute
leftover tests from a previous build. This change ensures CTestTestfile.cmake
gets regenerated so that ctest will see an empty test list in that case.
2024-11-27 11:02:54 +01:00
Ryan Houdek f41b9bc514 Merge pull request #4172 from alyssarosenzweig/jit/cf
OpcodeDispatcher: drop PossiblySetNZCV
2024-11-26 15:20:44 -08:00
Ryan Houdek 0463512c6c Merge pull request #4171 from Sonicadvance1/fix_ltrim
Utils/StringUtil: Fixes ltrim and adds a unittest
2024-11-26 13:09:22 -08:00
Ryan Houdek 47369d058e Utils/StringUtil: Fixes ltrim and adds a unittest
ltrim had the issue that it would always consume the left-most character
even if it wasn't whitespace. So `FEXLoader` would turn in to
`EXLoader`, even without any whitespace in the string.

Adds a test to ensure this doesn't occur again.
2024-11-26 12:57:48 -08:00
Billy Laws b6f34fa209 unittests: Add test for carry inversion bug 2024-11-26 09:02:23 -05:00
Alyssa Rosenzweig 03e0ca9833 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 09:02:01 -05:00
Alyssa Rosenzweig d19473160d OpcodeDispatcher: drop InvalidateDeferredFlags
it is now useless.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 09:02:01 -05:00
Alyssa Rosenzweig aeb2c98cbf OpcodeDispatcher: drop PossiblySetNZCV
this is a pain to track and, it turns out, buys us virtually nothing on flagm
systems. rip it out.

this fixes a bug with failing to set in all the right places.

on non-flagm systems there's a slight instcountci impact, but that is mostly
mitigated by the earlier patches in the series. so overall a wash there but
worth it for making the codebase easier
to reason about.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Alyssa Rosenzweig 2350ae5a07 OpcodeDispatcher: optimize BTC on !flagm
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Alyssa Rosenzweig 6076d1747e OpcodeDispatcher: use NZV invalidate CF set for BT
this is similar perf on flagm and better on not flagm.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Alyssa Rosenzweig f4ce6fb621 OpcodeDispatcher: optimize AAS/AAD flag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Ryan Houdek f0d25d413b InstcounCI: Update 2024-11-24 00:38:54 -08:00
Ryan Houdek c84503c271 unittests/ASM: Adds x87 loadstore tests for edge of pages
These tests ensure that FEX's x87 80-bit loadstores don't read/write
past the end of the page.
2024-11-24 00:36:50 -08:00
Ryan Houdek 1465df874b OpcodeDispatcher: Fixes 80-bit loads
Ensures reads don't go past the end of the page boundary.
SVE masked loads can make this more effective but `VLoadVectorMasked`
isn't setup to be efficient for this case yet.
2024-11-24 00:36:50 -08:00
Ryan Houdek 60c52e3826 Merge pull request #4167 from pmatos/CheckLDPathNonEmpty
Check that LDPath is not empty
2024-11-22 15:47:42 -08:00
Paulo Matos 13b806130b Check that LDPath is not empty
This ensures we don't underflow at --RootFSLength.
2024-11-22 10:24:14 +01:00
Ryan Houdek 22058c06a1 Merge pull request #4164 from bylaws/swwow
WOW64: Set the software CPU area flag
2024-11-19 09:39:33 -08:00
Ryan Houdek d5c96555f1 Merge pull request #4159 from asahilina/fexserver-two-sockets
FEXServer: Listen on both abstract & named sockets
2024-11-19 09:28:48 -08:00
Asahi Lina c3e8cd8d30 docs: Document that std::filesystem::temp_directory_path() is unsafe 2024-11-20 02:05:32 +09:00
Asahi Lina d761fc44f4 FEXServer: Listen on both abstract & named sockets
Abstract sockets have one limitation: they are bound to a network
namespace. Chromium/CEF sandboxes using a new netns, which breaks
connecting to the FEXServer.

To work around this, use and try *both* abstract and named sockets. As
long as either the filesystem or the network is unsandboxed, things will
work. If both are sandboxed, there isn't much we can do... but at that
point we shouldn't be reinitializing the FEXServer connection anyway
since the FS should be available on FEXInterpreter startup.
2024-11-20 01:58:17 +09:00
Asahi Lina a1aa2547ce FEXServerClient: Do not use strerror() in ConnectToServer()
This triggers glibc allocation.

Signed-off-by: Asahi Lina <lina@asahilina.net>
2024-11-20 01:58:17 +09:00
Asahi Lina 44213c3968 FEXServerClient: Switch GetTempFolder to not use temp_directory_path()
Apparently this causes allocations which are banned in some paths?
2024-11-19 01:46:33 +09:00
Billy Laws 830bd347c5 WOW64: Set the software CPU area flag
This is required for wow64.dll to setup the cross-process queue
that is used to pass through e.g. memory unmap events.
2024-11-18 16:18:31 +00:00
Ryan Houdek bcfdf39d63 Merge pull request #4157 from asahilina/fix-chromium-sandbox
Support CLONE_FS and CLONE_FILES with fork() semantics
2024-11-18 06:42:15 -08:00
Asahi Lina bfed21870f Support CLONE_FS and CLONE_FILES with fork() semantics
Needed by Discord, part of the Chromium sandbox code. The warning still
triggers because Chromium asks for CLONE_VM on x86_64, but that can be
safely ignored (CLONE_FS is the one that matters).
2024-11-18 22:09:33 +09:00
Ryan Houdek 4278c48791 Merge pull request #4158 from asahilina/hide-rootfs-fd-take2
FileManagement: Hide the FEX RootFS fd from /proc/self/fd take 2
2024-11-17 20:27:59 -08:00
Ryan Houdek c1db7a78b1 Merge pull request #4153 from pmatos/VixlSkip
Add some more tests as unsupported by vixl
2024-11-17 19:56:56 -08:00
Ryan Houdek 1bf06f8946 Merge pull request #4162 from pmatos/WarningAvoid
Avoid warning on assertionless builds
2024-11-17 19:03:31 -08:00
Ryan Houdek 09bfe58827 Merge pull request #4160 from asahilina/align-stack
FEXLoader: Align stack base
2024-11-17 18:27:25 -08:00
Paulo Matos 474c780399 Avoid warning on assertionless builds
This was causing unused variable warning due
to the variable only being used in an assertion.
2024-11-13 15:17:36 +01:00
Asahi Lina 4a67893f1d FEXLoader: Align stack base
This ensures that __libc_stack_end is aligned, the same way it is on
native.
2024-11-13 03:47:43 +09:00
Asahi Lina 73ffaa1e18 FileManagement: Hide the FEX RootFS fd from /proc/self/fd take 2
Apparently Chromium/CEF can chroot or otherwise sandbox the filesystem
away before forking and checking for directory FDs, making /proc
inaccessible, which means we can't stat it for our inode check, breaking
the hiding.

So, double down on things and do what Chromium does: open an fd to /proc
ahead of time, so that continues to work. Then we use it to update the
inode of our RootFS fd instead, and finally, also do the /proc fd itself
to hide that one too.

We also don't need to check the st_dev of /proc more than once, since
that's not expected to change anyway.

Fixes cefsimple.
2024-11-13 01:26:42 +09:00
Ryan Houdek e675f4241a Merge pull request #4154 from Liamolucko/check-home
Check if a candidate home directory exists before using it
2024-11-05 20:07:30 -08:00
Liam Murphy 7a61d9d2b4 Check if a candidate home directory exists before using it
This allows FEX to be used in situations where `HOME` is set to
something invalid, e.g. inside Nix builds.
2024-11-06 11:16:53 +11:00
Paulo Matos 06b950a9cd Rounding test doesn't need to be skipped 2024-11-04 19:02:12 +01:00
Paulo Matos de70651406 Add some more tests as unsupported by vixl
It seems a form of `mrs` is unsupported as well as the hint `wfe` used for `pause`.
Remove skipping Rounding(Neg|Pos).asm as they are passing.
2024-11-04 18:52:48 +01:00
LC 5ad7fdb2f3 Merge pull request #4149 from Sonicadvance1/iropsize_convert_class
IR: Convert OpSize over to enum class
2024-10-30 23:55:55 -04:00
Ryan Houdek 9b6cc8f7e0 IR: Convert OpSize over to enum class
NFC

Do the final mopping up to convert the OpSize enum to an enum class!
2024-10-29 16:52:16 -07:00
LC 5c6de4ed14 Merge pull request #4147 from Sonicadvance1/iropsize_convert_irops
IR: Converts base IR operations to store OpSize sizes
2024-10-29 12:02:11 -04:00
Ryan Houdek 82f936cb6d IR: Converts base IR operations to store OpSize sizes
NFC

Finally converts the IR operations themselves to store the OpSize for
the IR operation size and element sizes.

This also finally, FINALLY, converts that remaining `_Constant` helper
to stop using a size field that is specified in bits rather than bytes
like all the other IR op handlers. That thing was so confusing and now
it's gone.
2024-10-28 21:26:59 -07:00
LC 493b952e3f Merge pull request #4146 from Sonicadvance1/iropsize_implicit_jit
JIT: Remove implicit OpSize conversions
2024-10-28 23:47:36 -04:00
Ryan Houdek 460a21625e JIT: Remove implicit OpSize conversions
NFC
2024-10-28 19:48:40 -07:00
LC 00ab3f8440 Merge pull request #4145 from Sonicadvance1/iropsize_various_implicit
OpcodeDispatcher: Various missed OpSize implicit cast fixes
2024-10-28 22:38:41 -04:00
Ryan Houdek 034b62292b Passes/x86StackOptimization: Fixes implicit conversion of OpSize 2024-10-28 19:26:02 -07:00
Ryan Houdek 7b615a07d0 OpcodeDispatcher: Various missed OpSize implicit cast fixes
NFC
Probably more of these around, just tracking the few I found.
2024-10-28 19:18:36 -07:00
LC 704841f004 Merge pull request #4144 from Sonicadvance1/iropsize_addrsize
OpcodeDispatcher: Convert address size helpers to use OpSize
2024-10-28 22:12:40 -04:00
LC c122f3faf9 Merge pull request #4143 from Sonicadvance1/iropsize_flags
OpcodeDispatcher: Convert flags helpers over to OpSize
2024-10-28 22:10:19 -04:00
Ryan Houdek f74f276d64 OpcodeDispatcher: Convert address size helpers to use OpSize
NFC
2024-10-28 19:02:34 -07:00
Ryan Houdek 4b10cbdafd OpcodeDispatcher: Convert flags helpers over to OpSize
NFC

Plus the tertiary bits that require changing to support it.
2024-10-28 18:56:42 -07:00
LC 04c701e912 Merge pull request #4142 from Sonicadvance1/irsize_loadstoregpr
OpcodeDispatcher: Convert {Load,Store}GPRRegister to OpSize
2024-10-28 21:04:01 -04:00
LC 0c29f8faad Merge pull request #4141 from Sonicadvance1/missing_ir_sizes
IR: Fix some missing OpSize conversions
2024-10-28 21:02:46 -04:00
Ryan Houdek 5ed82fa0f6 OpcodeDispatcher: Convert {Load,Store}GPRRegister to OpSize
Trivial but quite a few places pass in a raw integer

NFC
2024-10-28 16:35:23 -07:00
Ryan Houdek 6cca007817 IR: Fix some missing OpSize conversions
Missed these in the previous PR.
2024-10-28 16:25:36 -07:00
LC c0a9463700 Merge pull request #4140 from Sonicadvance1/enforce_irsize
Convert all of the IR operations to use OpSize
2024-10-28 18:20:02 -04:00
Ryan Houdek 55b3d67eb4 Merge pull request #4138 from asahilina/hide-rootfs-fd
FileManagement: Hide the FEX RootFS fd from /proc/self/fd
2024-10-28 14:24:55 -07:00
Ryan Houdek 65ddae1b71 IR: Change F80VBSLStack to use IR::OpSize 2024-10-28 02:25:17 -07:00
Ryan Houdek 063f524084 IR: Change F80CVTToInt to use IR::OpSize 2024-10-28 02:24:25 -07:00
Ryan Houdek 2dd0a82059 IR: Change F80CVTTo to use IR::OpSize 2024-10-28 02:23:47 -07:00
Ryan Houdek b810070e9f IR: Change F80CVTInt to use IR::OpSize 2024-10-28 02:07:20 -07:00
Ryan Houdek 4c7ac17f7d IR: Change F80CVT to use IR::OpSize 2024-10-28 02:06:48 -07:00
Ryan Houdek 84767c8b20 IR: Change PushStack to use IR::OpSize 2024-10-28 02:05:35 -07:00
Ryan Houdek eccfb53bd5 IR: Change StoreStackMemory to use IR::OpSize 2024-10-28 02:01:47 -07:00
Ryan Houdek f4e930262f IR: Change PCLMUL to use IR::OpSize 2024-10-28 01:50:24 -07:00
Ryan Houdek d26d9e7e03 IR: Change CRC32 to use IR::OpSize 2024-10-28 01:50:24 -07:00
Ryan Houdek 51c1998d70 IR: Change VAESDecLast to use IR::OpSize 2024-10-28 01:50:24 -07:00
Ryan Houdek 87f818249d IR: Change VAESDec to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek d81f92f5e2 IR: Change VAESEncLast to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 29ffe02afe IR: Change VAESEnc to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek dfe4076fe4 IR: Change Vector_F64ToI32 to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 44f9df062e IR: Change Vector_FToI to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 6f98ef8cbb IR: Change VFCVTN2 to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek d54888a4c6 IR: Change VFCVTL2 to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 38c58706da IR: Change Vector_FToF to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 8ad9286bd4 IR: Change Vector_FToZS to use IR::OpSize 2024-10-28 01:29:32 -07:00
Ryan Houdek 03bc962564 IR: Change Vector_FToS to use IR::OpSize 2024-10-28 01:28:56 -07:00
Ryan Houdek 947b7ae6fe IR: Change Vector_SToF to use IR::OpSize 2024-10-28 01:23:20 -07:00
Ryan Houdek 30317ac979 IR: Change Float_FToF to use IR::OpSize 2024-10-28 01:22:45 -07:00
Ryan Houdek 4544e7c1af IR: Change Float_FromGPR_S to use IR::OpSize 2024-10-28 01:21:48 -07:00
Ryan Houdek 34050431ff IR: Change VDupFromGPR to use IR::OpSize 2024-10-28 01:20:06 -07:00
Ryan Houdek 65439956bf IR: Change VCastFromGPR to use IR::OpSize 2024-10-28 01:18:21 -07:00
Ryan Houdek a6cbce4fd7 IR: Change VFNMLS to use IR::OpSize 2024-10-28 01:16:11 -07:00
Ryan Houdek 17b851d4f3 IR: Change VFNMLA to use IR::OpSize 2024-10-28 01:15:47 -07:00
Ryan Houdek 362b5728be IR: Change VFMLS to use IR::OpSize 2024-10-28 01:15:25 -07:00
Ryan Houdek bd5159c7d5 IR: Change VFMLA to use IR::OpSize 2024-10-28 01:15:01 -07:00
Ryan Houdek 660dfcd1f9 IR: Change VFCADD to use IR::OpSize 2024-10-28 01:14:21 -07:00
Ryan Houdek 4e21177988 IR: Change VBSL to use IR::OpSize 2024-10-28 01:13:53 -07:00
Ryan Houdek 5a83a65905 IR: Change VTBX1 to use IR::OpSize 2024-10-28 01:13:15 -07:00
Ryan Houdek 061fc44923 IR: Change VTBL2 to use IR::OpSize 2024-10-28 01:13:00 -07:00
Ryan Houdek 0c7afa0672 IR: Change VTBL1 to use IR::OpSize 2024-10-28 01:12:43 -07:00
Ryan Houdek 30bf0d5767 IR: Change VFCMPUNO to use IR::OpSize 2024-10-28 01:12:22 -07:00
Ryan Houdek ee8e3127d2 IR: Change VFCMPORD to use IR::OpSize 2024-10-28 01:11:54 -07:00
Ryan Houdek e8f64f2976 IR: Change VFCMPLE to use IR::OpSize 2024-10-28 01:11:31 -07:00
Ryan Houdek 0a4b21da87 IR: Change VFCMPGT to use IR::OpSize 2024-10-28 01:11:11 -07:00
Ryan Houdek 074777bc75 IR: Change VFCMPLT to use IR::OpSize 2024-10-28 01:10:50 -07:00
Ryan Houdek 47c403998a IR: Change VFCMPNEQ to use IR::OpSize 2024-10-28 01:10:28 -07:00
Ryan Houdek 0fa095cec5 IR: Change VFCMPEQ to use IR::OpSize 2024-10-28 01:10:07 -07:00
Ryan Houdek cfa4e7f165 IR: Change VCMPGT to use IR::OpSize 2024-10-28 01:09:23 -07:00
Ryan Houdek 75b8226e7d IR: Change VCMPEQ to use IR::OpSize 2024-10-28 01:09:04 -07:00
Ryan Houdek 9e3c50ca2c IR: Change VExtr to use IR::OpSize 2024-10-28 01:08:35 -07:00
Ryan Houdek 365d8b9508 IR: Change VInsGPR to use IR::OpSize 2024-10-28 01:07:16 -07:00
Ryan Houdek 02ebe06496 IR: Change VInsElement to use IR::OpSize 2024-10-28 01:05:34 -07:00
Ryan Houdek 34d5e70e6b IR: Change VUShlSWide to use IR::OpSize 2024-10-28 00:59:33 -07:00
Ryan Houdek 2d8bd7b59d IR: Change VSShrSWide to use IR::OpSize 2024-10-28 00:57:02 -07:00
Ryan Houdek 3a6f5e638b IR: Change VUShrSWide to use IR::OpSize 2024-10-28 00:54:45 -07:00
Ryan Houdek c06274066f IR: Change VSShrS to use IR::OpSize 2024-10-28 00:51:19 -07:00
Ryan Houdek 820b0be9f2 IR: Change VUShrS to use IR::OpSize 2024-10-28 00:50:57 -07:00
Ryan Houdek 3e74817dd9 IR: Change VUShlS to use IR::OpSize 2024-10-28 00:50:36 -07:00
Ryan Houdek 3b9a2d2141 IR: Change VSShr to use IR::OpSize 2024-10-28 00:50:15 -07:00
Ryan Houdek 9ba431a51d IR: Change VUShr to use IR::OpSize 2024-10-28 00:49:52 -07:00
Ryan Houdek d0d2229db6 IR: Change VUShl to use IR::OpSize 2024-10-28 00:49:26 -07:00
Ryan Houdek 286258a2f2 IR: Change VUABDL2 to use IR::OpSize 2024-10-28 00:48:59 -07:00
Ryan Houdek ac14a88647 IR: Change VUABDL to use IR::OpSize 2024-10-28 00:48:32 -07:00
Ryan Houdek bf19578673 IR: Change VSMulH to use IR::OpSize 2024-10-28 00:47:44 -07:00
Ryan Houdek c5f396d889 IR: Change VUMulH to use IR::OpSize 2024-10-28 00:47:08 -07:00
Ryan Houdek 22325500d9 IR: Change VSMull2 to use IR::OpSize 2024-10-28 00:46:12 -07:00
Ryan Houdek 51eede080c IR: Change VUMull2 to use IR::OpSize 2024-10-28 00:45:50 -07:00
Ryan Houdek 503c86d47d IR: Change VSMull to use IR::OpSize 2024-10-28 00:45:23 -07:00
Ryan Houdek 0e0887181d IR: Change VUMull to use IR::OpSize 2024-10-28 00:44:20 -07:00
Ryan Houdek 974cc591bd IR: Change VMul to use IR::OpSize 2024-10-28 00:43:57 -07:00
Ryan Houdek 844afdb653 IR: Change VFMax to use IR::OpSize 2024-10-28 00:43:36 -07:00
Ryan Houdek c3c643d9b7 IR: Change VFMin to use IR::OpSize 2024-10-28 00:43:14 -07:00
Ryan Houdek b51b0c5f62 IR: Change VFDiv to use IR::OpSize 2024-10-28 00:42:54 -07:00
Ryan Houdek 96bdefddc9 IR: Change VFMul to use IR::OpSize 2024-10-28 00:42:32 -07:00
Ryan Houdek ba74a6a252 IR: Change VFSub to use IR::OpSize 2024-10-28 00:42:09 -07:00
Ryan Houdek 79427097e1 IR: Change VFAddV to use IR::OpSize 2024-10-28 00:41:46 -07:00
Ryan Houdek 7ec21b7121 IR: Change VFAddP to use IR::OpSize 2024-10-28 00:41:18 -07:00
Ryan Houdek 29115b3185 IR: Change VFAdd to use IR::OpSize 2024-10-28 00:38:55 -07:00
Ryan Houdek 33b3814642 IR: Change VTrn2 to use IR::OpSize 2024-10-28 00:37:44 -07:00
Ryan Houdek 85431a8132 IR: Change VTrn to use IR::OpSize 2024-10-28 00:35:12 -07:00
Ryan Houdek 2d9ef56a8b IR: Change VUnZip2 to use IR::OpSize 2024-10-28 00:31:48 -07:00
Ryan Houdek b2ae829731 IR: Change VUnZip to use IR::OpSize 2024-10-28 00:31:23 -07:00
Ryan Houdek a3c544a9a1 IR: Change VZip2 to use IR::OpSize 2024-10-28 00:27:28 -07:00
Ryan Houdek 9c2292289b IR: Change VZip to use IR::OpSize 2024-10-28 00:26:34 -07:00
Ryan Houdek b514548ca2 IR: Change VSMax to use IR::OpSize 2024-10-28 00:22:02 -07:00
Ryan Houdek 692a4a8fcd IR: Change VSMin to use IR::OpSize 2024-10-28 00:21:39 -07:00
Ryan Houdek 3a07cf7d70 IR: Change VUMax to use IR::OpSize 2024-10-28 00:21:22 -07:00
Ryan Houdek 7f5421fc26 IR: Change VUMin to use IR::OpSize 2024-10-28 00:21:00 -07:00
Ryan Houdek 04d62cd269 IR: Change VURAvg to use IR::OpSize 2024-10-28 00:20:35 -07:00
Ryan Houdek ac55e468a7 IR: Change VAddP to use IR::OpSize 2024-10-28 00:20:07 -07:00
Ryan Houdek e20db7dc88 IR: Change VSQSub to use IR::OpSize 2024-10-28 00:18:03 -07:00
Ryan Houdek 235bee9191 IR: Change VSQAdd to use IR::OpSize 2024-10-28 00:17:20 -07:00
Ryan Houdek 7a429b01c7 IR: Change VUQSub to use IR::OpSize 2024-10-28 00:16:31 -07:00
Ryan Houdek fbf5b14933 IR: Change VUQAdd to use IR::OpSize 2024-10-28 00:16:01 -07:00
Ryan Houdek 0dfd5dd96f IR: Change VXor to use IR::OpSize 2024-10-28 00:15:33 -07:00
Ryan Houdek 6d5acec958 IR: Change VOr to use IR::OpSize 2024-10-28 00:13:06 -07:00
Ryan Houdek beed43e577 IR: Change VAndn to use IR::OpSize 2024-10-28 00:12:37 -07:00
Ryan Houdek fc04b9113e IR: Change VAnd to use IR::OpSize 2024-10-28 00:11:44 -07:00
Ryan Houdek 68c038085a IR: Change VSub to use IR::OpSize 2024-10-28 00:10:21 -07:00
Ryan Houdek 176f5a2860 IR: Change VAdd to use IR::OpSize 2024-10-28 00:07:30 -07:00
Ryan Houdek e7c9623aa9 IR: Change VRev64 to use IR::OpSize 2024-10-27 23:36:03 -07:00
Ryan Houdek 53ca2ac378 IR: Change VRev32 to use IR::OpSize 2024-10-27 23:34:51 -07:00
Ryan Houdek d140cb4450 IR: Change VSQSHL to use IR::OpSize 2024-10-27 23:34:22 -07:00
Ryan Houdek 0ba501636e IR: Change VSRSHR to use IR::OpSize 2024-10-27 23:33:49 -07:00
Ryan Houdek ceca9fff17 IR: Change VSQXTUNPair to use IR::OpSize 2024-10-27 23:30:02 -07:00
Ryan Houdek 3e31abb645 IR: Change VSQXTUN2 to use IR::OpSize 2024-10-27 23:27:53 -07:00
Ryan Houdek 6c07cd319b IR: Change VSQXTUN to use IR::OpSize 2024-10-27 23:27:27 -07:00
Ryan Houdek 5247b7124f IR: Change VSQXTNPair to use IR::OpSize 2024-10-27 23:26:55 -07:00
Ryan Houdek 6cec557855 IR: Change VSQXTN2 to use IR::OpSize 2024-10-27 23:24:50 -07:00
Ryan Houdek 1868bd6777 IR: Change VSQXTN to use IR::OpSize 2024-10-27 23:24:25 -07:00
Ryan Houdek 7c7efeda82 IR: Change VUXTL2 to use IR::OpSize 2024-10-27 23:23:58 -07:00
Ryan Houdek 414486f1dd IR: Change VUXTL to use IR::OpSize 2024-10-27 23:23:12 -07:00
Ryan Houdek b8bc9659d4 IR: Change VSSHLL2 to use IR::OpSize 2024-10-27 23:22:47 -07:00
Ryan Houdek d52a6e6fc4 IR: Change VSSHLL to use IR::OpSize 2024-10-27 23:22:19 -07:00
Ryan Houdek 5ab41056ab IR: Change VSXTL2 to use IR::OpSize 2024-10-27 23:21:49 -07:00
Ryan Houdek 0692b34192 IR: Change VSXTL to use IR::OpSize 2024-10-27 23:21:17 -07:00
Ryan Houdek 3636c332ff IR: Change VUShrNI2 to use IR::OpSize 2024-10-27 23:05:41 -07:00
Ryan Houdek cb18963ded IR: Change VUShrNI to use IR::OpSize 2024-10-27 23:05:08 -07:00
Ryan Houdek 8cf92d3303 IR: Change VSShrI to use IR::OpSize 2024-10-27 23:03:21 -07:00
Ryan Houdek cdc5c15b4b IR: Change VUShraI to use IR::OpSize 2024-10-27 22:58:35 -07:00
Ryan Houdek ed313edd07 IR: Change VUShrI to use IR::OpSize 2024-10-27 22:57:57 -07:00
Ryan Houdek 9b981a4f61 IR: Change VShlI to use IR::OpSize 2024-10-27 22:55:08 -07:00
Ryan Houdek c791893b4a IR: Change VDupElement to use IR::OpSize 2024-10-27 22:50:47 -07:00
Ryan Houdek 2605c7e0b3 IR: Change VCMPLTZ to use IR::OpSize 2024-10-27 22:44:45 -07:00
Ryan Houdek f315948028 IR: Change VCMPGTZ to use IR::OpSize 2024-10-27 22:44:14 -07:00
Ryan Houdek 869367f7e2 IR: Change VCMPEQZ to use IR::OpSize 2024-10-27 22:43:42 -07:00
Ryan Houdek 3a2c7e8edd IR: Change VFRSqrt to use IR::OpSize 2024-10-27 22:43:12 -07:00
Ryan Houdek 0a34a43976 IR: Change VFSqrt to use IR::OpSize 2024-10-27 22:42:37 -07:00
Ryan Houdek 78cd21d78f IR: Change VFRecp to use IR::OpSize 2024-10-27 22:33:53 -07:00
Ryan Houdek 9f18de0196 IR: Change VFNeg to use IR::OpSize 2024-10-27 22:33:25 -07:00
Ryan Houdek 0bffdc4e27 IR: Change VFAbs to use IR::OpSize 2024-10-27 22:32:37 -07:00
Ryan Houdek 0f5ff53386 IR: Change VUMaxV to use IR::OpSize 2024-10-27 22:31:31 -07:00
Ryan Houdek 21611fc1ad IR: Change VUMinV to use IR::OpSize 2024-10-27 22:31:04 -07:00
Ryan Houdek 99e1eb5452 IR: Change VAddv to use IR::OpSize 2024-10-27 22:30:30 -07:00
Ryan Houdek 8c3ca44c57 IR: Change VPopcount to use IR::OpSize 2024-10-27 22:28:51 -07:00
Ryan Houdek 7a85e17d14 IR: Change VAbs to use IR::OpSize 2024-10-27 22:28:24 -07:00
Ryan Houdek b533dcd86d IR: Change VNot to use IR::OpSize 2024-10-27 22:27:28 -07:00
Ryan Houdek a379ce6fed IR: Change VNeg to use IR::OpSize 2024-10-27 22:09:02 -07:00
Ryan Houdek efd5c51110 IR: Change LoadNamedVectorIndexedConstant to use IR::OpSize 2024-10-27 22:08:27 -07:00
Ryan Houdek 886db4ffca IR: Change LoadNamedVectorConstant to use IR::OpSize 2024-10-27 22:05:42 -07:00
Ryan Houdek e6f6ee2bcd IR: Change VectorImm to use IR::OpSize 2024-10-27 21:58:34 -07:00
Ryan Houdek 4a6b5d4ec7 IR: Change VMov to use IR::OpSize 2024-10-27 21:52:11 -07:00
Ryan Houdek 027e7624cb IR: Change VFNMLSScalarInsert to use IR::OpSize 2024-10-27 18:37:30 -07:00
Ryan Houdek 0e31077735 IR: Change VFNMLAScalarInsert to use IR::OpSize 2024-10-27 18:36:53 -07:00
Ryan Houdek c8a9dd0d0a IR: Change VFMLSScalarInsert to use IR::OpSize 2024-10-27 18:36:14 -07:00
Ryan Houdek 2bd7ddaa31 IR: Change VFMLAScalarInsert to use IR::OpSize 2024-10-27 18:35:40 -07:00
Ryan Houdek 5566b4455b IR: Change VFCMPScalarInsert to use IR::OpSize 2024-10-27 18:35:03 -07:00
Ryan Houdek fed2c13521 IR: Change VFToIScalarInsert to use IR::OpSize 2024-10-27 18:32:11 -07:00
Ryan Houdek 37d092aab8 IR: Change VSToFGPRInsert to use IR::OpSize 2024-10-27 18:29:34 -07:00
Ryan Houdek 5626f4e50a IR: Change VSToFVectorInsert to use IR::OpSize 2024-10-27 18:27:39 -07:00
Ryan Houdek efbc42dac3 IR: Change VFToFScalarInsert to use IR::OpSize 2024-10-27 18:25:13 -07:00
Ryan Houdek 160934884d IR: Change VFRecpScalarInsert to use IR::OpSize 2024-10-27 18:21:05 -07:00
Ryan Houdek af1cfcb9bd IR: Change VFRSqrtScalarInsert to use IR::OpSize 2024-10-27 18:20:35 -07:00
Ryan Houdek d6f726fc23 IR: Change VFSqrtScalarInsert to use IR::OpSize 2024-10-27 18:20:06 -07:00
Ryan Houdek 92ee071eb2 IR: Change VFMaxScalarInsert to use IR::OpSize 2024-10-27 18:17:43 -07:00
Ryan Houdek bdfa8ad4f3 IR: Change VFMinScalarInsert to use IR::OpSize 2024-10-27 18:17:17 -07:00
Ryan Houdek 37540f4927 IR: Change VFDivScalarInsert to use IR::OpSize 2024-10-27 18:16:30 -07:00
Ryan Houdek 000ab5ff19 IR: Change VFMulScalarInsert to use IR::OpSize 2024-10-27 18:15:58 -07:00
Ryan Houdek f054274948 IR: Change VFSubScalarInsert to use IR::OpSize 2024-10-27 18:15:22 -07:00
Ryan Houdek 081907e168 IR: Change VFAddScalarInsert to use IR::OpSize 2024-10-27 18:14:37 -07:00
Ryan Houdek 1a115a8ce6 IR: Change FCmp to use IR::OpSize 2024-10-27 18:06:34 -07:00
Ryan Houdek fd9158c75f IR: Change Float_ToGPR_ZS to use IR::OpSize 2024-10-27 18:03:18 -07:00
Ryan Houdek 1bde30a196 IR: Change Float_ToGPR_S to use IR::OpSize 2024-10-27 18:02:45 -07:00
Ryan Houdek 764aacaa8f IR: Change VExtractToGPR to use IR::OpSize 2024-10-27 17:58:53 -07:00
Ryan Houdek a848211926 IR: Change NZCVSelectV to use IR::OpSize 2024-10-27 17:53:15 -07:00
Ryan Houdek f1a42869d5 IR: Change CondJump to use IR::OpSize 2024-10-27 17:50:51 -07:00
Ryan Houdek 97a6ba9931 IR: Change VLoadNonTemporal to use IR::OpSize 2024-10-27 17:45:29 -07:00
Ryan Houdek f4744f1e79 IR: Change VStoreNonTemporalPair to use IR::OpSize 2024-10-27 17:44:57 -07:00
Ryan Houdek 321f686108 IR: Change VStoreNonTemporal to use IR::OpSize 2024-10-27 17:44:23 -07:00
Ryan Houdek 90340350fa IR: Change MemCpy to use IR::OpSize 2024-10-27 17:43:10 -07:00
Ryan Houdek 6f4fd4467b IR: Change MemSet to use IR::OpSize 2024-10-27 17:42:39 -07:00
Ryan Houdek b31ce13f68 IR: Change Pop to use IR::OpSize 2024-10-27 17:41:30 -07:00
Ryan Houdek 260d3b0b4e IR: Change Push to use IR::OpSize 2024-10-27 17:39:21 -07:00
Ryan Houdek c8c7ffbf05 IR: Change VBroadcastFromMem to use IR::OpSize 2024-10-27 17:35:56 -07:00
Ryan Houdek 4b03185b77 IR: Change VStoreVectorElement to use IR::OpSize 2024-10-27 17:32:25 -07:00
Ryan Houdek 52ec572db3 IR: Change VLoadVectorElement to use IR::OpSize 2024-10-27 17:27:42 -07:00
Ryan Houdek dc31cf83c6 IR: Change VLoadVectorGatherMaskedQPS to use IR::OpSize 2024-10-27 17:22:07 -07:00
Ryan Houdek 8a4f51257d IR: Change VLoadVectorGatherMasked to use IR::OpSize 2024-10-27 17:21:33 -07:00
Ryan Houdek 051469fa16 IR: Change VStoreVectorMasked to use IR::OpSize 2024-10-27 17:20:56 -07:00
Ryan Houdek 3f6cdc2e03 IR: Change VLoadVectorMasked to use IR::OpSize 2024-10-27 17:20:18 -07:00
Ryan Houdek f3449f2b00 IR: Change StoreMemTSO to use IR::OpSize 2024-10-27 17:17:26 -07:00
Ryan Houdek d7691d9a25 IR: Change LoadMemTSO to use IR::OpSize 2024-10-27 17:16:23 -07:00
Ryan Houdek f414d4934c IR: Change StoreMemPair to use IR::OpSize 2024-10-27 17:15:00 -07:00
Ryan Houdek 014917301a IR: Change StoreMem to use IR::OpSize 2024-10-27 17:14:18 -07:00
Ryan Houdek cc483acbde IR: Change LoadMemPair to use IR::OpSize 2024-10-27 16:33:52 -07:00
Ryan Houdek 07f8a4eadd IR: Change LoadMem to use IR::OpSize 2024-10-27 16:33:14 -07:00
Ryan Houdek 5fd127b53a IR: Change StoreContextIndexed to use IR::OpSize 2024-10-27 15:51:30 -07:00
Ryan Houdek ece89ddeab IR: Change LoadContextIndexed to use IR::OpSize 2024-10-27 15:50:03 -07:00
Ryan Houdek a1565a7d99 IR: Change StoreContextPair to use IR::OpSize 2024-10-27 15:47:08 -07:00
Ryan Houdek 2f9b0de742 IR: Change StoreContext to use IR::OpSize 2024-10-27 15:46:32 -07:00
Ryan Houdek 7e5f1b5859 IR: Change LoadContextPair to use IR::OpSize 2024-10-27 15:42:39 -07:00
Ryan Houdek 40fd4bbb66 IR: Change LoadContext to use IR::OpSize 2024-10-27 15:42:07 -07:00
Ryan Houdek e4143352c9 IR: Change Store{PF,AF} to use IR::OpSize 2024-10-27 15:37:41 -07:00
Ryan Houdek c045e14837 IR: Change StoreRegister to use IR::OpSize 2024-10-27 15:37:05 -07:00
Ryan Houdek f0f3c215ce IR: Change Load{PF,AF} to use IR::OpSize 2024-10-27 15:35:26 -07:00
Ryan Houdek 8f4113d859 IR: Change LoadRegister to use IR::OpSize 2024-10-27 15:34:46 -07:00
Ryan Houdek 4cfc2ac1a4 IR: Change AllocateFPR to use IR::OpSize 2024-10-27 15:30:11 -07:00
LC d2aa5217dc Merge pull request #4132 from Sonicadvance1/fix_irsize
Fix IR operation usage to use OpSize when possible
2024-10-27 17:26:11 -04:00
Ryan Houdek e8baf4a28c OpcodeDispatcher: Ensure IR ops use OpSize
NFC
2024-10-27 14:11:25 -07:00
Ryan Houdek e438d32879 OpcodeDispatcher/Vector: Ensure IR ops use OpSize
NFC
2024-10-27 14:11:07 -07:00
Ryan Houdek 32ef10b273 OpcodeDispatcher/AVX128: Ensure IR ops use OpSize
NFC
2024-10-27 14:11:07 -07:00
Ryan Houdek ad296051b7 OpcodeDispatcher/Crypto: Ensure IR ops use OpSize
NFC
2024-10-27 14:11:07 -07:00
Ryan Houdek e603136918 OpcodeDispatcher/Flags: Ensure IR ops use OpSize
NFC
2024-10-27 14:11:07 -07:00
Ryan Houdek f8a61f7d7e OpcodeDispatcher/X87: Ensure IR ops use OpSize
NFC
2024-10-27 14:11:07 -07:00
Ryan Houdek cb5ba8baae OpcodeDispatcher/X87F64: Ensure IR ops use OpSize
NFC
2024-10-27 14:11:07 -07:00
LC 079e70fc4e Merge pull request #4134 from Sonicadvance1/move_jit
JIT: Moves Arm64 JIT up one folder
2024-10-27 17:10:26 -04:00
Asahi Lina 3d701f5fcf FileManagement: Hide the FEX RootFS fd from /proc/self/fd
Chromium/CEF has code that iterates through all open FDs and bails if
any are directories (apparently a sandboxing sanity check). To avoid
this check, we need to hide the RootFS FD. This requires hooking all the
getdents variants to skip that entry.

To keep the runtime cost low, we keep track of the inode of
/proc/self/fd/<rootfs fd> (note: not the RootFS inode, the inode of the
magic symlink in /proc), and first do a quick check on that. If it
matches, then we stat the dirfd we are reading and check against the
procfs device, to complete the inode equality check.

As an extra benefit, this also fixes code that tries to iterate and
close all/extra FDs and ends up closing the RootFS fd.
2024-10-27 07:05:00 +09:00
Ryan Houdek b31e4a3c27 JIT: Moves Arm64 JIT up one folder
We have only one JIT, there is no reason to subfolder this. Move it up
one folder.

NFC
2024-10-25 18:44:20 -07:00
LC 992d6e8477 Merge pull request #4136 from Sonicadvance1/support_tpidrro
FEXCore: Adds support for CPU Index through TPIDRRO
2024-10-25 21:43:33 -04:00
LC f6cdb165a3 Merge pull request #4137 from Sonicadvance1/minor_pushf_opt
OpcodeDispatcher: Minor optimization to small pushf
2024-10-25 21:42:45 -04:00
LC f60388d160 Merge pull request #4135 from Sonicadvance1/remove_xop
X86Tables: Removes XOP tables
2024-10-25 19:41:31 -04:00
Ryan Houdek 96fa2ad8eb InstcountCI: Update 2024-10-25 15:43:10 -07:00
Ryan Houdek f143462ebe OpcodeDispatcher: Minor optimization to small pushf
The push operation already truncates the result, there's no need to bfe
it. Noticed this while cleaning up in #4134. Removes one instruction for
16-bit and 32-bit pushf instructions.
2024-10-25 15:41:22 -07:00
Ryan Houdek 51fa61a1cd InstcountCI: Adds missing pushf implementations
pushf was aliasing to pushfq, needed an o16 prefix.
We also weren't testing the 32-bit path, which only exists on 32-bit, so
add that as well.
2024-10-25 15:40:44 -07:00
Ryan Houdek 048e967546 FEXCore: Adds support for CPU Index through TPIDRRO 2024-10-25 15:07:57 -07:00
Ryan Houdek 608fd49ac3 CodeEmitter/unittests: Add support for TPIDRRO_EL0 2024-10-25 15:07:29 -07:00
Ryan Houdek bb630797b5 CodeEmitter: Add support for TPIDRRO_EL0 2024-10-25 15:07:15 -07:00
Ryan Houdek 01a6e914f2 X86Tables: Removes XOP tables
These weren't even wired up to the frontend. We aren't going to support
XOP, so just remove the tables.
2024-10-25 14:32:13 -07:00
Ryan Houdek 0190e1a00b Merge pull request #4130 from pmatos/X87F64Simp
X87 Code Simplification
2024-10-24 10:57:12 -07:00
Paulo Matos 11a87c22f9 instcountci: X87 code simplification 2024-10-24 18:17:48 +02:00
Paulo Matos 5f6c0d2245 X87 code simplification
Merges some of the code from reduced precision into the main path
since they are practically the same.
2024-10-24 18:15:49 +02:00
Ryan Houdek caaacb6c15 Merge pull request #4127 from alyssarosenzweig/opt/masking
Optimize bsf, bsr, register cmpxchg, pcmpistri
2024-10-23 07:28:29 -07:00
Ryan Houdek 767c61c08b Merge pull request #4129 from pmatos/RPRESInstcountci
Disable RPRES in instcounci files
2024-10-23 07:27:57 -07:00
Paulo Matos dc93e30451 Disable RPRES in instcounci files
This was giving false changes on RPRES enabled HW.
2024-10-23 15:09:17 +02:00
Ryan Houdek 368162df87 Merge pull request #4128 from ahoneybun/update-ubuntu-support
add Ubuntu 24.10 and remove unsupported releases
2024-10-22 14:43:22 -07:00
Aaron Honeycutt 9eb2106ed2 update supported list 2024-10-22 15:35:41 -06:00
Aaron Honeycutt cfc05b78fe add Ubuntu 24.10 2024-10-22 15:27:21 -06:00
Alyssa Rosenzweig d2a42c0038 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:39:21 -04:00
Alyssa Rosenzweig 58a3d174ec OpcodeDispatcher: explain why we provide defined bsf behaviour
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:36:24 -04:00
Alyssa Rosenzweig 9c605e7333 OpcodeDispatcher: optimize bsf/bsr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:31:49 -04:00
Alyssa Rosenzweig 1fd7e88ffd OpcodeDispatcher: optimize cmp in cmpxchg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:29:47 -04:00
Alyssa Rosenzweig c5e7da0631 OpcodeDispatcher: optimize more cmpxchg mask
none of it matters.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:29:47 -04:00
Alyssa Rosenzweig 68f58e415f OpcodeDispatcher: optimize cmpxchg masking
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:29:47 -04:00
Alyssa Rosenzweig 698abec25c JIT: drop FindMSB zero handling
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:29:47 -04:00
Alyssa Rosenzweig 1578f5ed47 JIT: drop FindLSB masking
consequence of the UB

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:29:47 -04:00
Alyssa Rosenzweig 80d7b5a5c9 JIT: drop FindLSB zero handling
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:29:47 -04:00
Alyssa Rosenzweig eb023ceb51 IR: make FindLSB/FindMSB undefined for zero
these are used in places:

* bsf/bsr
* pcmpblabla
* x87 fild

In all cases we explicitly check for zero and change the behaviour accordingly.
So weaken the IR op to let us optimize. No sense checking twice.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-22 13:29:47 -04:00
LC 5d1fda7d7f Merge pull request #4125 from Sonicadvance1/unify_pshuflhw
OpcodeDispatcher: Unify PSHUF{L,H}W implementations
2024-10-21 14:37:39 -04:00
Ryan Houdek fafc04a59e OpcodeDispatcher/AVX128: Fixes glibc allocation 2024-10-21 09:34:59 -07:00
Ryan Houdek ddcca58f64 InstcountCI: Update for PSHUF{L,H}W changes 2024-10-21 09:25:58 -07:00
Ryan Houdek fe8f5c745d OpcodeDispatcher/AVX128: Use unified PSHUF{L,H}W implementation
Allows the AVX128 implementation to use the same implementation as the
128-bit SSE implementation, because it works per 128-bit lane just like
SSE. Is a minor optimization.
2024-10-21 09:24:38 -07:00
Ryan Houdek d876224358 OpcodeDispatcher: Unify MMX and SSE PSHUF{L,H}W implementations
No functional change.
2024-10-21 09:11:05 -07:00
LC c4306f2f0a Merge pull request #4122 from Sonicadvance1/avx_fixes
AVX128: Fixes some AVX bugs
2024-10-18 19:34:28 -04:00
Ryan Houdek 8b7a227820 InstcountCI: Update 2024-10-18 15:49:25 -07:00
Ryan Houdek d7afcee622 AVX128: Fixes asome AVX bugs
vblendvps, vblendvpd, vpblendvb all broke because of failing to zext the
128-bit register correctly.

vinsertps broke because it was accidentally using the wrong
implementation.

Extends each of their unittests to handle these cases.
2024-10-18 15:49:25 -07:00
Ryan Houdek a421ff1105 Merge pull request #4093 from pmatos/FXtractFix
Fix FXTRACT for 0.0 and -0.0
2024-10-17 05:31:34 -07:00
Paulo Matos 5997030c97 instcountci: Fix FXTRACT for 0.0 and -0.0 2024-10-17 09:13:40 +02:00
Paulo Matos 3c8086373b FXTRACT fix ASM tests 2024-10-17 09:05:17 +02:00
Paulo Matos 10ec6b63b6 Fix FXTRACT for 0.0 and -0.0
Fixes fxtract by returning the correct values for 0.0 and -0.0. We moved the split of fxtract into _sig and _exp, to the opcode dispatcher, to ease some comparisons.

Also removed the IR node F80XTRACTStack which is not needed anymore.
2024-10-17 09:05:10 +02:00
Paulo Matos 49087007be Implement NZCVSelectV for selection on FPRs
Behaves like NZCVSelect for FPRs.
2024-10-17 08:42:50 +02:00
Ryan Houdek 0897cd8777 Merge pull request #4116 from Sonicadvance1/personality_handling
LinuxEmulation: Personality handling
2024-10-16 15:52:35 -07:00
Ryan Houdek 4b945a9041 Merge pull request #4105 from pmatos/X87MMXState
Implement explicit state switch between X87 and MMX
2024-10-16 15:52:27 -07:00
Ryan Houdek d66ed71bc6 Merge pull request #4082 from Sonicadvance1/fix_wine
FEXLoader: Fixes newer wine versions and Fedora
2024-10-16 15:52:19 -07:00
LC b5b34df155 Merge pull request #4119 from Sonicadvance1/unaligned_lock
unittests/ASM: Adds missing unaligned atomic tests
2024-10-15 12:16:37 -04:00
LC ff51435747 Merge pull request #4118 from Sonicadvance1/wfe_for_pause
FEXCore: Change yield implementation to use wfe
2024-10-15 12:15:11 -04:00
Paulo Matos def561986b instcountci: Implements explicit state switch between X87 and MMX 2024-10-15 17:58:59 +02:00
Paulo Matos 5c258d4a2a ASM Test: Implements explicit state switch between X87 and MMX
Tags is set to all valid in FEX, but in host it's set to all valid _and_
reinterpreted. Adding this to known failures in the host runner.
2024-10-15 17:58:59 +02:00
Paulo Matos 0d53f2b45c Implements explicit state switch between X87 and MMX
Fixes #3850
2024-10-15 17:58:53 +02:00
Ryan Houdek 4f03044fe7 unittests/ASM: Adds missing unaligned atomic tests
Fixes #2670

Walked through all the unaligned atomic tests to find which ones were
missing. Turns out it was only ADC, NEG, NOT, and SBB.
2024-10-15 06:59:38 -07:00
Ryan Houdek b967538435 InstcountCI: Add pause instruction 2024-10-15 05:52:41 -07:00
Ryan Houdek e53f3969e9 FEXCore: Change yield implementation to use wfe
According to
https://github.com/rust-lang/rust/commit/c064b6560b7ce0adeb9bbf5d7dcf12b1acb0c807
turns out that the arm yield instruction is effectively a nop on all
reasonably new CPUs.
Instead switch over to wfe because it matches x86 `PAUSE` semantics more
closely.
2024-10-15 05:50:26 -07:00
Ryan Houdek 5026bf8247 Merge pull request #4117 from pmatos/NoTest
Remove file since FXAM_Simple is not a Linux test
2024-10-11 18:17:09 -07:00
Paulo Matos 09cb4f5fc5 Remove file since FXAM_Simple is not a Linux test
Already properly skipped in the right place.
File added accidentally.
2024-10-11 16:55:07 +02:00
Paulo Matos ddd7a550e4 Remove check on top 16bits
Makes it uniform among all 3DNow tests instead of
some checking and some don't.
2024-10-11 16:18:43 +02:00
Ryan Houdek 3398f22c16 FEXLinuxTests: Adds personality test
These would have failed before the prior changes.
2024-10-11 05:03:28 -07:00
Ryan Houdek 7f17519fbf LinuxEmulation/personality: Support PER_LINUX32 2024-10-11 04:52:23 -07:00
Ryan Houdek a65884f9ae LinuxEmulation/personality: Support UNAME26 2024-10-11 04:52:21 -07:00
Ryan Houdek d70766f4c8 LinuxEmulation: Support personality tracking
Doesn't handle the emulation of it, but handle passing it to the host
kernel, tracking the value, and inheriting it through new threads.
2024-10-11 04:52:21 -07:00
Ryan Houdek 1365aa8881 FEXLinuxTests: Update to c++20 2024-10-11 04:52:21 -07:00
Ryan Houdek fe5bc02682 FEXLoader: Fixes newer wine versions and Fedora
This was brought up by #3831 but I finally got the courage to look at
the hard problem.

Although I'm only tackling half of the problem with this PR, which is
that FEXLoader needs to strip the rootfs path from the executed path if
it begins with the rootfs, plus some changes to the surrounding code.

The primary concern here is that when an application has been executed
under FEX, specifically through binfmt_misc, then FEX needs to prepend
the full rootfs path otherwise Linux can't find the program.
Additionally execveat with an FD will resolve a full path to the rootfs.

So past FEX's initial setup, we need to strip off the rootfs path to
provide an "absolute" path that is visible to the guest application
later. Which is kind of funny since we have a `RootFSRedirect` function
which did the exact opposite. This was due to legacy problems in the
original ELFLoader that couldn't handle symlinks correctly, which has
since been resolved, so that no longer needs to exist.

There was also some weirdness in `GetApplicationNames` where the passed
in argument list was modifying Args[0] and then saving the Program as
well. Which I just got rid of. Also stopped passing in the arguments by
value because....why did I write it like that?

In InterpreterHandler we now need to check if we can open the path
inside the rootfs or fallback without it. Plus I had to change the
shebang handling so it stopped prefixing the rootfs AGAIN. Took the time
to change the shebang handling there so it stops creating string copies
and instead just generates views.

Overall this fixes a fairly major flaw with how we were representing
`/proc/self` to the application, which was breaking wine since it would
prefix the rootfs multiple times, which was weird.

It doesn't address the remaining problem in #3831, which is that
applications can still see some of the leaky abstractions with symlinks
through the rootfs, but I want to get at least this step in.
2024-10-11 01:40:36 -07:00
Ryan Houdek 6f096e7c4b FHU: Add StringArgumentParser function
Split this out so we can unittest it.

Adds a unittest to handle specific edge cases.
2024-10-11 01:40:36 -07:00
LC 389ad737e6 Merge pull request #4103 from Sonicadvance1/shared_vdso_mmap
VDSOEmulation: Support loading VDSO thunk as shared
2024-10-10 01:25:12 -04:00
LC c00f7813a2 Merge pull request #4108 from Sonicadvance1/ensure_x87_size_save_restore
unittests/ASM: Ensures FNSAVE and FRSTOR only store as much data as required
2024-10-10 01:00:54 -04:00
LC e5ceaa182d Merge pull request #4110 from Sonicadvance1/remove_unused_memory_regions
unitests/ASM: Removes unused MemoryRegion configs
2024-10-10 00:59:31 -04:00
LC eeb8eb1824 Merge pull request #4109 from Sonicadvance1/fix_fsgs_testharness
TestHarnessRunner: Fixes FS/GS usage in tests
2024-10-10 00:56:20 -04:00
Ryan Houdek 7c6444c37c unitests/ASM: Removes unused MemoryRegion configs
FEX's ASM unitests had the problem that they were copy and pasted
templates and MemoryRegion was copied in to almost all tests.

Very few tests actually use the MemoryRegion they were asking for and
instead used none, or the hardcoded memory regions that the
TestHarnessRunner provides.

This is entirely a sed replacement and minor fixups plus reverts for the
few tests that actually use the region asked for.
2024-10-08 16:07:25 -07:00
Ryan Houdek 4a179c8f87 TestHarnessRunner: Fixes FS/GS usage in tests
When writing `FEX_bugs/tls_vector_element.asm` I had to switch to using
GS segment instead of FS segment because the unittests didn't correctly
restore FS after running. This is because GS is unused on Linux
applications, but FS would become broken and break glibc cleanup on
shutdown.

Now that xbyak has been updated to v7.09, it now supports
{rd,wr}{fs,gs}base which allows us to save and restore the segments
correctly. This lets us drop in TLS tests in to unittests more easily
now.

Ensured this works by modifying the test temporarily to use fs instead
of gs again, seeing it crash before HostRunner changes, and work after
HostRunner changes.

Fixes #4104
2024-10-08 15:42:44 -07:00
Ryan Houdek 2b3895a514 Update xbyak to v7.09 2024-10-08 15:38:00 -07:00
Ryan Houdek 0bb0f9cec7 unittests/ASM: Ensures FNSAVE and FRSTOR only store as much data as required
The instruction definition only allows these instructions to load/store
94 or 108 bytes, not affecting any bytes afterwards. This is a bit
awkward because 80-bit x87 registers are stored at the end.

FEX has an optimization today where it uses overlapping loads and stores
for the first seven x87 registers, and a split loadstore for the final
register. This ensures that we get the correct data while reducing the
number of loadstores.

We didn't have a unittest in place to ensure we only ever write the
correct amount of data, so changes like in #4107 which look correct from
an initial glance, would have resulted in broken behaviour.

This unittest ensures both that the instructions don't try to access
beyond the end of the page, and also ensures that they don't overwrite
subsequent data. Making sure that potentially broken behaviour doesn't
make its way in.
2024-10-08 15:31:36 -07:00
Ryan Houdek 6a07ea73a8 Merge pull request #4106 from slp/compat-input-prctl
FEXCore: adds support for compat input prctl
2024-10-08 13:22:19 -07:00
Sergio Lopez 5c51c54ccc FEXCore: adds support for compat input prctl
The size of the input_event struct differs between 32 bits applications
and 64 bits applications. To deal with this, the kernel implements a
compat variant for the input syscalls, but it's only enabled for 32 bit
processes.

In libkrunfw we're introducing a prctl that enables a 64 bit process to
request the kernel to enable the compat variant for the input syscalls.
This commit makes use of that interface for enabling/disabling the
compat input variant as required.

The visible effect is that input devices such as gamepads work properly
on emulated 32 bit applications.

Signed-off-by: Sergio Lopez <slp@redhat.com>
2024-10-08 11:38:42 +02:00
Ryan Houdek c740801ea5 VDSOEmulation: Support loading VDSO thunk as shared
Just requires a thread pointer check to be fixed in FEXCore.
This doesn't need unique pages to exist for the file mapping and can be
shared since it's readonly mapped.
2024-10-03 21:06:35 -07:00
1320 changed files with 102252 additions and 75920 deletions

No files matched your search

+4 -1
View File
@@ -3,10 +3,13 @@
# Ignore all files in the External directory
External/*
# SoftFloat-3e code doesn't belong to us
# SoftFloat-3e code doesn't belong to us
FEXCore/Source/Common/SoftFloat-3e/*
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
FEXCore/Source/Interface/Core/X86Tables/*
# Inline headers with list-like content that can't be processed individually
Source/Tools/LinuxEmulation/LinuxSyscalls/x*/SyscallsNames.inl
Source/Tools/LinuxEmulation/LinuxSyscalls/x*/Ioctl/*.inl
+3
View File
@@ -13,3 +13,6 @@
# Second reformat to find fixed point PR#3577
905aa935f5ce344a48ef4d5edab3c31efa8d793e
# Reformat of CodeEmitter inl files
8760c593ece92d7e9fa94c40da0368fd367c9cad
+1 -1
View File
@@ -250,7 +250,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
+1 -1
View File
@@ -184,7 +184,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
+1 -1
View File
@@ -97,7 +97,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
+2 -2
View File
@@ -128,7 +128,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
@@ -137,7 +137,7 @@ jobs:
- name: Upload results InstCountCI
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}-instcountci
+1 -1
View File
@@ -92,7 +92,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
+1 -1
View File
@@ -126,7 +126,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
+3
View File
@@ -47,3 +47,6 @@
[submodule "External/jemalloc_glibc"]
path = External/jemalloc_glibc
url = https://github.com/FEX-Emu/jemalloc.git
[submodule "External/tracy"]
path = External/tracy
url = https://github.com/wolfpld/tracy
+33 -13
View File
@@ -8,7 +8,7 @@ option(BUILD_TESTS "Build unit tests to ensure sanity" TRUE)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(BUILD_FEXCONFIG "Build FEXConfig" TRUE)
option(ENABLE_CLANG_THUNKS "Build thunks with clang" FALSE)
option(ENABLE_CLANG_THUNKS "Build thunks with clang" TRUE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
@@ -26,17 +26,17 @@ option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_VIXL_SIMULATOR "Forces the FEX JIT to use the VIXL simulator" FALSE)
option(ENABLE_VIXL_SIMULATOR "Enable use of VIXL simulator for emulation (only useful for CI testing)" FALSE)
option(ENABLE_VIXL_DISASSEMBLER "Enables debug disassembler output with VIXL" FALSE)
option(USE_LEGACY_BINFMTMISC "Uses legacy method of setting up binfmt_misc" FALSE)
option(COMPILE_VIXL_DISASSEMBLER "Compiles the vixl disassembler in to vixl" FALSE)
option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling capabilities" FALSE)
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend you want to use for the FEXCore profiler")
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use for the FEXCore profiler (gpuvis, tracy)")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_DEV_ROOTFS "/" CACHE FILEPATH "Path to the sysroot used for cross-compiling for i686 and x86_64")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
@@ -61,6 +61,22 @@ if (ENABLE_FEXCORE_PROFILER)
if (FEXCORE_PROFILER_BACKEND STREQUAL "GPUVIS")
add_definitions(-DFEXCORE_PROFILER_BACKEND=1)
elseif (FEXCORE_PROFILER_BACKEND STREQUAL "TRACY")
add_definitions(-DFEXCORE_PROFILER_BACKEND=2)
add_definitions(-DTRACY_ENABLE=1)
# Required so that Tracy will only start in the selected guest application
add_definitions(-DTRACY_MANUAL_LIFETIME=1)
add_definitions(-DTRACY_DELAYED_INIT=1)
# This interferes with FEX's signal handling
add_definitions(-DTRACY_NO_CRASH_HANDLER=1)
# Tracy can gather call stack samples in regular intervals, but this
# isn't useful for us since it would usually sample opaque JIT code
add_definitions(-DTRACY_NO_SAMPLING=1)
# This pulls in libbacktrace which allocators in global constructors (before FEX can set up its allocator hooks)
add_definitions(-DTRACY_NO_CALLSTACK=1)
if (MINGW_BUILD)
message(FATAL_ERROR "Tracy profiler not supported")
endif()
else()
message(FATAL_ERROR "Unknown FEXCore profiler backend ${FEXCORE_PROFILER_BACKEND}")
endif()
@@ -265,16 +281,15 @@ set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-poin
include_directories(External/robin-map/include/)
if (BUILD_TESTS)
# Enable vixl disassembler if tests are enabled.
set(COMPILE_VIXL_DISASSEMBLER TRUE)
endif()
if (COMPILE_VIXL_DISASSEMBLER OR ENABLE_VIXL_SIMULATOR)
if (BUILD_TESTS OR ENABLE_VIXL_DISASSEMBLER OR ENABLE_VIXL_SIMULATOR)
add_subdirectory(External/vixl/)
include_directories(SYSTEM External/vixl/src/)
endif()
if (ENABLE_FEXCORE_PROFILER AND FEXCORE_PROFILER_BACKEND STREQUAL "TRACY")
add_subdirectory(External/tracy)
endif()
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# This means we were attempted to get compiled with GCC
message(FATAL_ERROR "FEX doesn't support getting compiled with GCC!")
@@ -298,7 +313,7 @@ add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
if (BUILD_TESTS)
find_package(Catch2 QUIET)
find_package(Catch2 3 QUIET)
if (NOT Catch2_FOUND)
add_subdirectory(External/Catch2/)
@@ -412,10 +427,13 @@ configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/include/Config.h.in
${CMAKE_BINARY_DIR}/generated/ConfigDefines.h)
include(CTest)
if (BUILD_TESTS)
include(CTest)
enable_testing()
message(STATUS "Unit tests are enabled")
if (NOT BUILD_TESTING)
# CMake checks this variable before generating CTestTestfile.cmake
message(SEND_ERROR "Unit tests require BUILD_TESTING to be enabled")
endif()
set (TEST_JOB_COUNT "" CACHE STRING "Override number of parallel jobs to use while running tests")
if (TEST_JOB_COUNT)
@@ -476,6 +494,7 @@ if (BUILD_THUNKS)
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
"-DX86_DEV_ROOTFS=${X86_DEV_ROOTFS}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
@@ -494,6 +513,7 @@ if (BUILD_THUNKS)
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
"-DX86_DEV_ROOTFS=${X86_DEV_ROOTFS}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
+157 -183
View File
@@ -11,6 +11,14 @@
* FEX-Emu ALU operations usually have a 32-bit or 64-bit operating size encoded in the IR operation,
* This allows FEX to use a single helper function which decodes to both handlers.
*/
#pragma once
#ifndef INCLUDED_BY_EMITTER
#include <CodeEmitter/Emitter.h>
namespace ARMEmitter {
struct EmitterOps : Emitter {
#endif
private:
static bool IsADRRange(int64_t Imm) {
return Imm >= -1048576 && Imm <= 1048575;
@@ -28,26 +36,23 @@ public:
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adr(ARMEmitter::Register rd, BackwardLabel const* Label) {
void adr(ARMEmitter::Register rd, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void adr(ARMEmitter::Register rd, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::ADR });
void adr(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::ADR});
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
}
void adr(ARMEmitter::Register rd, BiDirectionalLabel *Label) {
void adr(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
adr(rd, &Label->Backward);
}
else {
} else {
adr(rd, &Label->Forward);
}
}
@@ -57,39 +62,34 @@ public:
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adrp(ARMEmitter::Register rd, BackwardLabel const* Label) {
void adrp(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void adrp(ARMEmitter::Register rd, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::ADRP });
void adrp(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::ADRP});
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
}
void adrp(ARMEmitter::Register rd, BiDirectionalLabel *Label) {
void adrp(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
adrp(rd, &Label->Backward);
}
else {
} else {
adrp(rd, &Label->Forward);
}
}
void LongAddressGen(ARMEmitter::Register rd, BackwardLabel const* Label) {
void LongAddressGen(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>());
if (IsADRRange(Imm)) {
// If the range is in ADR range then we can just use ADR.
adr(rd, Label);
}
else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL)
- (GetCursorAddress<int64_t>() & ~0xFFFLL);
} else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
// If the range is in the ADRP range then we can use ADRP.
bool NeedsOffset = !IsADRPAligned(reinterpret_cast<uint64_t>(Label->Location));
@@ -102,24 +102,22 @@ public:
// Now even an add
add(ARMEmitter::Size::i64Bit, rd, rd, AlignedOffset);
}
}
else {
} else {
LOGMAN_MSG_A_FMT("Unscaled offset too large");
FEX_UNREACHABLE;
}
}
void LongAddressGen(ARMEmitter::Register rd, ForwardLabel* Label) {
Label->Insts.emplace_back(SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::LONG_ADDRESS_GEN });
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::LONG_ADDRESS_GEN});
// Emit a register index and a nop. These will be backpatched.
dc32(rd.Idx());
nop();
}
void LongAddressGen(ARMEmitter::Register rd, BiDirectionalLabel *Label) {
void LongAddressGen(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
LongAddressGen(rd, &Label->Backward);
}
else {
} else {
LongAddressGen(rd, &Label->Forward);
}
}
@@ -176,11 +174,7 @@ public:
// Logical immediate
void and_(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
and_(s, rd, rn, n, immr, imms);
}
@@ -191,11 +185,7 @@ public:
void ands(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
ands(s, rd, rn, n, immr, imms);
}
@@ -206,22 +196,14 @@ public:
void orr(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
orr(s, rd, rn, n, immr, imms);
}
void eor(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
eor(s, rd, rn, n, immr, imms);
}
@@ -355,8 +337,8 @@ public:
const auto lsb_p_width = lsb + width;
LOGMAN_THROW_A_FMT(width >= 1, "bfxil needs width >= 1");
LOGMAN_THROW_A_FMT(lsb_p_width <= reg_size_bits, "bfxil lsb + width ({}) must be <= {}. lsb={}, width={}",
lsb_p_width, reg_size_bits, lsb, width);
LOGMAN_THROW_A_FMT(lsb_p_width <= reg_size_bits, "bfxil lsb + width ({}) must be <= {}. lsb={}, width={}", lsb_p_width, reg_size_bits,
lsb, width);
bfm(s, rd, rn, lsb, lsb_p_width - 1);
}
@@ -375,188 +357,142 @@ public:
// Data processing - 2 source
void udiv(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0000'10U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0000'10U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void sdiv(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0000'11U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0000'11U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void lslv(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'00U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0010'00U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void lsrv(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'01U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0010'01U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void asrv(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'10U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0010'10U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void rorv(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'11U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0010'11U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void crc32b(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'00U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0100'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32h(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'01U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0100'01U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32w(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'10U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0100'10U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32cb(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'00U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0101'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32ch(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'01U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0101'01U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32cw(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'10U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0101'10U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void smax(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0110'00U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0110'00U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void umax(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0110'01U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0110'01U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void smin(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0110'10U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0110'10U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void umin(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0110'11U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0110'11U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void subp(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0000'00U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0000'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void irg(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0001'00U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0001'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void gmi(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0001'01U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0001'01U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void pacga(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0011'00U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0011'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void crc32x(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'11U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0100'11U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void crc32cx(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'11U << 10);
constexpr uint32_t Op = (0b001'1010'110U << 21) | (0b0101'11U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void subps(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b011'1010'110U << 21) |
(0b0000'00U << 10);
constexpr uint32_t Op = (0b011'1010'110U << 21) | (0b0000'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i64Bit, rd, rn, rm);
}
// Data processing - 1 source
void rbit(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'00U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0000'00U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void rev16(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'01U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0000'01U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void rev(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'10U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0000'10U << 10);
DataProcessing_1Source(Op, ARMEmitter::Size::i32Bit, rd, rn);
}
void rev32(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'10U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0000'10U << 10);
DataProcessing_1Source(Op, ARMEmitter::Size::i64Bit, rd, rn);
}
void clz(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0001'00U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0001'00U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void cls(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0001'01U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0001'01U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void rev(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'11U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0000'11U << 10);
DataProcessing_1Source(Op, ARMEmitter::Size::i64Bit, rd, rn);
}
void rev(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'10U << 10) |
(s == ARMEmitter::Size::i64Bit ? (1U << 10) : 0);
uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0000'10U << 10) | (s == ARMEmitter::Size::i64Bit ? (1U << 10) : 0);
DataProcessing_1Source(Op, s, rd, rn);
}
void ctz(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0001'10U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0001'10U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void cnt(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0001'11U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0001'11U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void abs(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0010'00U << 10);
constexpr uint32_t Op = (0b101'1010'110U << 21) | (0b0'0000U << 16) | (0b0010'00U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
@@ -573,27 +509,33 @@ public:
orr(ARMEmitter::Size::i32Bit, rd.R(), ARMEmitter::Reg::zr, rn.R(), ARMEmitter::ShiftType::LSL, 0);
}
void mvn(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void mvn(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL,
uint32_t amt = 0) {
orn(s, rd, ARMEmitter::Reg::zr, rn, Shift, amt);
}
void and_(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void and_(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b000'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void ands(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void ands(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b110'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void bic(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void bic(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b000'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void bics(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void bics(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b110'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void orr(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void orr(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b010'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
@@ -601,30 +543,36 @@ public:
ands(s, Reg::zr, rn, rm, shift, amt);
}
void orn(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void orn(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b010'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void eor(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void eor(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b100'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void eon(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void eon(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b100'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
// AddSub - shifted register
void add(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void add(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
add(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void adds(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void adds(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
adds(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void cmn(ARMEmitter::XRegister rn, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
adds(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::zr, rn.R(), rm.R(), Shift, amt);
}
void sub(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void sub(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
sub(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void neg(ARMEmitter::XRegister rd, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
@@ -633,23 +581,27 @@ public:
void cmp(ARMEmitter::XRegister rn, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, rn.R(), rm.R(), Shift, amt);
}
void subs(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void subs(ARMEmitter::XRegister rd, ARMEmitter::XRegister rn, ARMEmitter::XRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void negs(ARMEmitter::XRegister rd, ARMEmitter::XRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(rd, ARMEmitter::XReg::zr, rm, Shift, amt);
}
void add(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void add(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
add(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void adds(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void adds(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
adds(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void cmn(ARMEmitter::WRegister rn, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
adds(ARMEmitter::Size::i32Bit, ARMEmitter::WReg::zr, rn.R(), rm.R(), Shift, amt);
}
void sub(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void sub(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
sub(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void neg(ARMEmitter::WRegister rd, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
@@ -658,65 +610,78 @@ public:
void cmp(ARMEmitter::WRegister rn, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::rsp, rn.R(), rm.R(), Shift, amt);
}
void subs(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void subs(ARMEmitter::WRegister rd, ARMEmitter::WRegister rn, ARMEmitter::WRegister rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void negs(ARMEmitter::WRegister rd, ARMEmitter::WRegister rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(rd, ARMEmitter::WReg::zr, rm, Shift, amt);
}
void add(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
void add(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_A_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b000'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void adds(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
void adds(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_A_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b010'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void cmn(ARMEmitter::Size s, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void cmn(ARMEmitter::Size s, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL,
uint32_t amt = 0) {
adds(s, ARMEmitter::Reg::zr, rn, rm, Shift, amt);
}
void sub(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
void sub(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_A_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b100'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void neg(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void neg(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL,
uint32_t amt = 0) {
sub(s, rd, ARMEmitter::Reg::zr, rm, Shift, amt);
}
void cmp(ARMEmitter::Size s, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void cmp(ARMEmitter::Size s, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL,
uint32_t amt = 0) {
subs(s, ARMEmitter::Reg::zr, rn, rm, Shift, amt);
}
void subs(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
void subs(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_A_FMT(Shift != ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b110'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void negs(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
void negs(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift = ARMEmitter::ShiftType::LSL,
uint32_t amt = 0) {
subs(s, rd, ARMEmitter::Reg::zr, rm, Shift, amt);
}
// AddSub - extended register
void add(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
LOGMAN_THROW_AA_FMT(Shift <= 4, "Shift amount is too large");
void add(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option,
uint32_t Shift = 0) {
LOGMAN_THROW_A_FMT(Shift <= 4, "Shift amount is too large");
constexpr uint32_t Op = 0b000'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
void adds(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
void adds(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option,
uint32_t Shift = 0) {
constexpr uint32_t Op = 0b010'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
void cmn(ARMEmitter::Size s, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
adds(s, ARMEmitter::Reg::zr, rn, rm, Option, Shift);
}
void sub(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
void sub(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option,
uint32_t Shift = 0) {
constexpr uint32_t Op = 0b100'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
void subs(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
void subs(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option,
uint32_t Shift = 0) {
constexpr uint32_t Op = 0b110'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
@@ -751,8 +716,8 @@ public:
// Rotate right into flags
void rmif(XRegister rn, uint32_t shift, uint32_t mask) {
LOGMAN_THROW_AA_FMT(shift <= 63, "Shift must be within 0-63. Shift: {}", shift);
LOGMAN_THROW_AA_FMT(mask <= 15, "Mask must be within 0-15. Mask: {}", mask);
LOGMAN_THROW_A_FMT(shift <= 63, "Shift must be within 0-63. Shift: {}", shift);
LOGMAN_THROW_A_FMT(mask <= 15, "Mask must be within 0-15. Mask: {}", mask);
uint32_t Op = 0b1011'1010'0000'0000'0000'0100'0000'0000;
Op |= rn.Idx() << 5;
@@ -816,7 +781,8 @@ public:
}
void cset(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0001'1010'100 << 21;
ConditionalCompare(Op, 0, 0b01, s, rd, ARMEmitter::Reg::zr, ARMEmitter::Reg::zr, static_cast<ARMEmitter::Condition>(FEXCore::ToUnderlying(Cond) ^ FEXCore::ToUnderlying(ARMEmitter::Condition::CC_NE)));
ConditionalCompare(Op, 0, 0b01, s, rd, ARMEmitter::Reg::zr, ARMEmitter::Reg::zr,
static_cast<ARMEmitter::Condition>(FEXCore::ToUnderlying(Cond) ^ FEXCore::ToUnderlying(ARMEmitter::Condition::CC_NE)));
}
void csinc(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0001'1010'100 << 21;
@@ -898,8 +864,7 @@ public:
private:
static constexpr Condition InvertCondition(Condition cond) {
// These behave as always, so it makes no sense to allow inverting these.
LOGMAN_THROW_AA_FMT(cond != Condition::CC_AL && cond != Condition::CC_NV,
"Cannot invert CC_AL or CC_NV");
LOGMAN_THROW_A_FMT(cond != Condition::CC_AL && cond != Condition::CC_NV, "Cannot invert CC_AL or CC_NV");
return static_cast<Condition>(FEXCore::ToUnderlying(cond) ^ 1);
}
@@ -950,7 +915,7 @@ private:
LSL12 = true;
Imm >>= 12;
}
LOGMAN_THROW_AA_FMT(TooLarge == false, "Imm amount too large: 0x{:x}", Imm);
LOGMAN_THROW_A_FMT(TooLarge == false, "Imm amount too large: 0x{:x}", Imm);
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
@@ -995,7 +960,8 @@ private:
}
// Logical immediate
void DataProcessing_Logical_Imm(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint32_t n, uint32_t immr, uint32_t imms) {
void DataProcessing_Logical_Imm(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint32_t n,
uint32_t immr, uint32_t imms) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
@@ -1014,9 +980,8 @@ private:
[[maybe_unused]] const auto lsb_p_width = lsb + width;
const auto reg_size_bits = RegSizeInBits(s);
LOGMAN_THROW_AA_FMT(lsb_p_width <= reg_size_bits, "lsb + width ({}) must be <= {}. lsb={}, width={}",
lsb_p_width, reg_size_bits, lsb, width);
LOGMAN_THROW_AA_FMT(width >= 1, "xbfiz width must be >= 1");
LOGMAN_THROW_A_FMT(lsb_p_width <= reg_size_bits, "lsb + width ({}) must be <= {}. lsb={}, width={}", lsb_p_width, reg_size_bits, lsb, width);
LOGMAN_THROW_A_FMT(width >= 1, "xbfiz width must be >= 1");
const auto immr = (reg_size_bits - lsb) & (reg_size_bits - 1);
const auto imms = width - 1;
@@ -1028,12 +993,13 @@ private:
}
}
void DataProcessing_Extract(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, uint32_t Imm) {
void DataProcessing_Extract(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm,
uint32_t Imm) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
// Current ARMv8 spec hardcodes SF == N for this class of instructions.
// Anythign else is undefined behaviour.
const uint32_t N = s == ARMEmitter::Size::i64Bit ? (1U << 22) : 0;
const uint32_t N = s == ARMEmitter::Size::i64Bit ? (1U << 22) : 0;
uint32_t Instr = Op;
@@ -1076,10 +1042,11 @@ private:
}
// AddSub - shifted register
void DataProcessing_Shifted_Reg(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ShiftType Shift, uint32_t amt) {
LOGMAN_THROW_AA_FMT((amt & ~0b11'1111U) == 0, "Shift amount too large");
void DataProcessing_Shifted_Reg(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn,
ARMEmitter::Register rm, ARMEmitter::ShiftType Shift, uint32_t amt) {
LOGMAN_THROW_A_FMT((amt & ~0b11'1111U) == 0, "Shift amount too large");
if (s == ARMEmitter::Size::i32Bit) {
LOGMAN_THROW_AA_FMT(amt < 32, "Shift amount for 32-bit must be below 32");
LOGMAN_THROW_A_FMT(amt < 32, "Shift amount for 32-bit must be below 32");
}
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
@@ -1097,7 +1064,8 @@ private:
}
// AddSub - extended register
void DataProcessing_Extended_Reg(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::ExtendedType Option, uint32_t Shift) {
void DataProcessing_Extended_Reg(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn,
ARMEmitter::Register rm, ARMEmitter::ExtendedType Option, uint32_t Shift) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
@@ -1113,7 +1081,8 @@ private:
}
// Conditional compare - register
template<typename T>
void ConditionalCompare(uint32_t Op, uint32_t o1, uint32_t o2, uint32_t o3, ARMEmitter::Size s, ARMEmitter::Register rn, T rm, ARMEmitter::StatusFlags flags, ARMEmitter::Condition Cond) {
void ConditionalCompare(uint32_t Op, uint32_t o1, uint32_t o2, uint32_t o3, ARMEmitter::Size s, ARMEmitter::Register rn, T rm,
ARMEmitter::StatusFlags flags, ARMEmitter::Condition Cond) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
@@ -1131,7 +1100,8 @@ private:
}
template<typename T>
void ConditionalCompare(uint32_t Op, uint32_t o1, uint32_t o2, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, T rm, ARMEmitter::Condition Cond) {
void ConditionalCompare(uint32_t Op, uint32_t o1, uint32_t o2, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, T rm,
ARMEmitter::Condition Cond) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
@@ -1148,7 +1118,8 @@ private:
}
// Data-processing - 3 source
void DataProcessing_3Source(uint32_t Op, uint32_t Op0, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, ARMEmitter::Register rm, ARMEmitter::Register ra) {
void DataProcessing_3Source(uint32_t Op, uint32_t Op0, ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn,
ARMEmitter::Register rm, ARMEmitter::Register ra) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
@@ -1170,4 +1141,7 @@ private:
dc32(Instr);
}
#ifndef INCLUDED_BY_EMITTER
}; // struct LoadstoreEmitterOps
} // namespace ARMEmitter
#endif
File diff suppressed because it is too large. Load diff
+291 -305
View File
@@ -3,339 +3,325 @@
*
* Most of these instructions will use `BackwardLabel`, `ForwardLabel`, or `BiDirectionLabel` to determine where a branch targets.
*/
#pragma once
#ifndef INCLUDED_BY_EMITTER
#include <CodeEmitter/Emitter.h>
namespace ARMEmitter {
struct EmitterOps : Emitter {
#endif
public:
// Branches, Exception Generating and System instructions
public:
// Conditional branch immediate
///< Branch conditional
void b(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm);
public:
// Conditional branch immediate
///< Branch conditional
void b(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm);
}
void b(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
}
void b(ARMEmitter::Condition Cond, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, 0);
}
void b(ARMEmitter::Condition Cond, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
b(Cond, &Label->Backward);
} else {
b(Cond, &Label->Forward);
}
void b(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
}
///< Branch consistent conditional
void bc(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm);
}
void bc(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
}
void bc(ARMEmitter::Condition Cond, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, 0);
}
void bc(ARMEmitter::Condition Cond, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
bc(Cond, &Label->Backward);
} else {
bc(Cond, &Label->Forward);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void b(ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, 0);
}
// Unconditional branch register
void br(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 | 0b0'000 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
UnconditionalBranch(Op, rn);
}
void blr(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 | 0b0'001 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
UnconditionalBranch(Op, rn);
}
void ret(ARMEmitter::Register rn = ARMEmitter::Reg::r30) {
constexpr uint32_t Op = 0b1101011 << 25 | 0b0'010 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
UnconditionalBranch(Op, rn);
}
// Unconditional branch immediate
void b(uint32_t Imm) {
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm);
}
void b(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
}
void b(ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::B});
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, 0);
}
void b(BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
b(&Label->Backward);
} else {
b(&Label->Forward);
}
}
void b(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
b(Cond, &Label->Backward);
}
else {
b(Cond, &Label->Forward);
}
void bl(uint32_t Imm) {
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm);
}
void bl(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
}
void bl(ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::B});
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, 0);
}
void bl(BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
bl(&Label->Backward);
} else {
bl(&Label->Forward);
}
}
///< Branch consistent conditional
void bc(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm);
// Compare and branch
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, 0);
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
cbz(s, rt, &Label->Backward);
} else {
cbz(s, rt, &Label->Forward);
}
void bc(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, 0);
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
cbnz(s, rt, &Label->Backward);
} else {
cbnz(s, rt, &Label->Forward);
}
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void bc(ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, 0);
// Test and branch immediate
void tbz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::TEST_BRANCH});
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, 0);
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
tbz(rt, Bit, &Label->Backward);
} else {
tbz(rt, Bit, &Label->Forward);
}
}
void bc(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
bc(Cond, &Label->Backward);
}
else {
bc(Cond, &Label->Forward);
}
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0111 << 24;
// Unconditional branch register
void br(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'000 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
TestAndBranch(Op, rt, Bit, Imm);
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
UnconditionalBranch(Op, rn);
}
void blr(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'001 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
constexpr uint32_t Op = 0b0011'0111 << 24;
UnconditionalBranch(Op, rn);
}
void ret(ARMEmitter::Register rn = ARMEmitter::Reg::r30) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'010 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
UnconditionalBranch(Op, rn);
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::TEST_BRANCH});
constexpr uint32_t Op = 0b0011'0111 << 24;
// Unconditional branch immediate
void b(uint32_t Imm) {
constexpr uint32_t Op = 0b0001'01 << 26;
TestAndBranch(Op, rt, Bit, 0);
}
UnconditionalBranch(Op, Imm);
}
void b(BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void b(LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::B });
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, 0);
}
void b(BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
b(&Label->Backward);
}
else {
b(&Label->Forward);
}
}
void bl(uint32_t Imm) {
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm);
}
void bl(BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void bl(LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::B });
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, 0);
}
void bl(BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
bl(&Label->Backward);
}
else {
bl(&Label->Forward);
}
}
// Compare and branch
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, 0);
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbz(s, rt, &Label->Backward);
}
else {
cbz(s, rt, &Label->Forward);
}
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, 0);
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbnz(s, rt, &Label->Backward);
}
else {
cbnz(s, rt, &Label->Forward);
}
}
// Test and branch immediate
void tbz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, 0);
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbz(rt, Bit, &Label->Backward);
}
else {
tbz(rt, Bit, &Label->Forward);
}
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbnz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, 0);
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbnz(rt, Bit, &Label->Backward);
}
else {
tbnz(rt, Bit, &Label->Forward);
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
tbnz(rt, Bit, &Label->Backward);
} else {
tbnz(rt, Bit, &Label->Forward);
}
}
private:
// Conditional branch immediate
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, ARMEmitter::Condition Cond, uint32_t Imm) {
uint32_t Instr = Op;
// Conditional branch immediate
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, ARMEmitter::Condition Cond, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= Op1 << 24;
Instr |= (Imm & 0x7'FFFF) << 5;
Instr |= Op0 << 4;
Instr |= FEXCore::ToUnderlying(Cond);
Instr |= Op1 << 24;
Instr |= (Imm & 0x7'FFFF) << 5;
Instr |= Op0 << 4;
Instr |= FEXCore::ToUnderlying(Cond);
dc32(Instr);
}
dc32(Instr);
}
// Unconditional branch register
void UnconditionalBranch(uint32_t Op, ARMEmitter::Register rn) {
uint32_t Instr = Op;
Instr |= Encode_rn(rn);
dc32(Instr);
}
// Unconditional branch register
void UnconditionalBranch(uint32_t Op, ARMEmitter::Register rn) {
uint32_t Instr = Op;
Instr |= Encode_rn(rn);
dc32(Instr);
}
// Unconditional branch - immediate
void UnconditionalBranch(uint32_t Op, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= Imm & 0x3FF'FFFF;
dc32(Instr);
}
// Unconditional branch - immediate
void UnconditionalBranch(uint32_t Op, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= Imm & 0x3FF'FFFF;
dc32(Instr);
}
// Compare and branch
void CompareAndBranch(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
// Compare and branch
void CompareAndBranch(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
uint32_t Instr = Op;
Instr |= SF;
Instr |= (Imm & 0x7'FFFF) << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
Instr |= SF;
Instr |= (Imm & 0x7'FFFF) << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
// Test and branch - immediate
void TestAndBranch(uint32_t Op, ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
uint32_t Instr = Op;
// Test and branch - immediate
void TestAndBranch(uint32_t Op, ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= (Bit >> 5) << 31;
Instr |= (Bit & 0b1'1111) << 19;
Instr |= (Imm & 0x3FFF) << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
Instr |= (Bit >> 5) << 31;
Instr |= (Bit & 0b1'1111) << 19;
Instr |= (Imm & 0x3FFF) << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
#ifndef INCLUDED_BY_EMITTER
}; // struct LoadstoreEmitterOps
} // namespace ARMEmitter
#endif
-4
View File
@@ -95,10 +95,6 @@ public:
protected:
void ResetBuffer() {
CurrentOffset = BufferBase;
}
uint8_t* BufferBase;
uint8_t* CurrentOffset;
uint64_t Size;
+94 -94
View File
@@ -341,93 +341,88 @@ public:
};
template<uint32_t op0, uint32_t op1, uint32_t CRn, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenSystemReg() {
return op0 << 19 | op1 << 16 | CRn << 12 | CRm << 8 | op2 << 5;
};
inline constexpr uint32_t GenSystemReg = op0 << 19 | op1 << 16 | CRn << 12 | CRm << 8 | op2 << 5;
// This `SystemRegister` enum is used for the mrs/msr instructions.
enum class SystemRegister : uint32_t {
CTR_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b001>(),
DCZID_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b111>(),
TPIDR_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b010>(),
RNDR = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b000>(),
RNDRRS = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b001>(),
NZCV = GenSystemReg<0b11, 0b011, 0b0100, 0b0010, 0b000>(),
FPCR = GenSystemReg<0b11, 0b011, 0b0100, 0b0100, 0b000>(),
CNTFRQ_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b000>(),
CNTVCT_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b010>(),
CTR_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b001>,
DCZID_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b111>,
TPIDR_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b010>,
RNDR = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b000>,
RNDRRS = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b001>,
NZCV = GenSystemReg<0b11, 0b011, 0b0100, 0b0010, 0b000>,
FPCR = GenSystemReg<0b11, 0b011, 0b0100, 0b0100, 0b000>,
TPIDRRO_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b011>,
CNTFRQ_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b000>,
CNTVCT_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b010>,
};
template<uint32_t op1, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenDCReg() {
return op1 << 16 | CRm << 8 | op2 << 5;
};
inline constexpr uint32_t GenDCReg = op1 << 16 | CRm << 8 | op2 << 5;
// This `DataCacheOperation` enum is used for the dc instruction.
enum class DataCacheOperation : uint32_t {
IVAC = GenDCReg<0b000, 0b0110, 0b001>(),
ISW = GenDCReg<0b000, 0b0110, 0b010>(),
CSW = GenDCReg<0b000, 0b1010, 0b010>(),
CISW = GenDCReg<0b000, 0b1110, 0b010>(),
ZVA = GenDCReg<0b011, 0b0100, 0b001>(),
CVAC = GenDCReg<0b011, 0b1010, 0b001>(),
CVAU = GenDCReg<0b011, 0b1011, 0b001>(),
CIVAC = GenDCReg<0b011, 0b1110, 0b001>(),
IVAC = GenDCReg<0b000, 0b0110, 0b001>,
ISW = GenDCReg<0b000, 0b0110, 0b010>,
CSW = GenDCReg<0b000, 0b1010, 0b010>,
CISW = GenDCReg<0b000, 0b1110, 0b010>,
ZVA = GenDCReg<0b011, 0b0100, 0b001>,
CVAC = GenDCReg<0b011, 0b1010, 0b001>,
CVAU = GenDCReg<0b011, 0b1011, 0b001>,
CIVAC = GenDCReg<0b011, 0b1110, 0b001>,
// MTE2
IGVAC = GenDCReg<0b000, 0b0110, 0b011>(),
IGSW = GenDCReg<0b000, 0b0110, 0b100>(),
IGDVAC = GenDCReg<0b000, 0b0110, 0b101>(),
IGDSW = GenDCReg<0b000, 0b0110, 0b110>(),
CGSW = GenDCReg<0b000, 0b1010, 0b100>(),
CGDSW = GenDCReg<0b000, 0b1010, 0b110>(),
CIGSW = GenDCReg<0b000, 0b1110, 0b100>(),
CIGDSW = GenDCReg<0b000, 0b1110, 0b110>(),
IGVAC = GenDCReg<0b000, 0b0110, 0b011>,
IGSW = GenDCReg<0b000, 0b0110, 0b100>,
IGDVAC = GenDCReg<0b000, 0b0110, 0b101>,
IGDSW = GenDCReg<0b000, 0b0110, 0b110>,
CGSW = GenDCReg<0b000, 0b1010, 0b100>,
CGDSW = GenDCReg<0b000, 0b1010, 0b110>,
CIGSW = GenDCReg<0b000, 0b1110, 0b100>,
CIGDSW = GenDCReg<0b000, 0b1110, 0b110>,
// MTE
GVA = GenDCReg<0b011, 0b0100, 0b011>(),
GZVA = GenDCReg<0b011, 0b0100, 0b100>(),
CGVAC = GenDCReg<0b011, 0b1010, 0b011>(),
CGDVAC = GenDCReg<0b011, 0b1010, 0b101>(),
CGVAP = GenDCReg<0b011, 0b1100, 0b011>(),
CGDVAP = GenDCReg<0b011, 0b1100, 0b101>(),
CGVADP = GenDCReg<0b011, 0b1101, 0b011>(),
CGDVADP = GenDCReg<0b011, 0b1101, 0b101>(),
CIGVAC = GenDCReg<0b011, 0b1110, 0b011>(),
CIGDVAC = GenDCReg<0b011, 0b1110, 0b101>(),
GVA = GenDCReg<0b011, 0b0100, 0b011>,
GZVA = GenDCReg<0b011, 0b0100, 0b100>,
CGVAC = GenDCReg<0b011, 0b1010, 0b011>,
CGDVAC = GenDCReg<0b011, 0b1010, 0b101>,
CGVAP = GenDCReg<0b011, 0b1100, 0b011>,
CGDVAP = GenDCReg<0b011, 0b1100, 0b101>,
CGVADP = GenDCReg<0b011, 0b1101, 0b011>,
CGDVADP = GenDCReg<0b011, 0b1101, 0b101>,
CIGVAC = GenDCReg<0b011, 0b1110, 0b011>,
CIGDVAC = GenDCReg<0b011, 0b1110, 0b101>,
// DPB
CVAP = GenDCReg<0b011, 0b1100, 0b001>(),
CVAP = GenDCReg<0b011, 0b1100, 0b001>,
// DPB2
CVADP = GenDCReg<0b011, 0b1101, 0b001>(),
CVADP = GenDCReg<0b011, 0b1101, 0b001>,
};
template<uint32_t CRm, uint32_t op2>
constexpr uint32_t GenHintBarrierReg() {
return CRm << 8 | op2 << 5;
}
inline constexpr uint32_t GenHintBarrierReg = CRm << 8 | op2 << 5;
// This `HintRegister` enum is used for the hint instruction.
enum class HintRegister : uint32_t {
NOP = GenHintBarrierReg<0b0000, 0b000>(),
YIELD = GenHintBarrierReg<0b0000, 0b001>(),
WFE = GenHintBarrierReg<0b0000, 0b010>(),
WFI = GenHintBarrierReg<0b0000, 0b011>(),
SEV = GenHintBarrierReg<0b0000, 0b100>(),
SEVL = GenHintBarrierReg<0b0000, 0b101>(),
DGH = GenHintBarrierReg<0b0000, 0b110>(),
CSDB = GenHintBarrierReg<0b0010, 0b100>(),
NOP = GenHintBarrierReg<0b0000, 0b000>,
YIELD = GenHintBarrierReg<0b0000, 0b001>,
WFE = GenHintBarrierReg<0b0000, 0b010>,
WFI = GenHintBarrierReg<0b0000, 0b011>,
SEV = GenHintBarrierReg<0b0000, 0b100>,
SEVL = GenHintBarrierReg<0b0000, 0b101>,
DGH = GenHintBarrierReg<0b0000, 0b110>,
CSDB = GenHintBarrierReg<0b0010, 0b100>,
};
// This `BarrierRegister` enum is used for the various barrier instructions.
enum class BarrierRegister : uint32_t {
CLREX = GenHintBarrierReg<0b0000, 0b010>(),
TCOMMIT = GenHintBarrierReg<0b0000, 0b011>(),
DSB = GenHintBarrierReg<0b0000, 0b100>(),
DMB = GenHintBarrierReg<0b0000, 0b101>(),
ISB = GenHintBarrierReg<0b0000, 0b110>(),
SB = GenHintBarrierReg<0b0000, 0b111>(),
CLREX = GenHintBarrierReg<0b0000, 0b010>,
TCOMMIT = GenHintBarrierReg<0b0000, 0b011>,
DSB = GenHintBarrierReg<0b0000, 0b100>,
DMB = GenHintBarrierReg<0b0000, 0b101>,
ISB = GenHintBarrierReg<0b0000, 0b110>,
SB = GenHintBarrierReg<0b0000, 0b111>,
};
// This `BarrierScope` enum is used for the dsb/dmb instructions.
@@ -512,7 +507,7 @@ enum class SVEFMaxMinImm : uint32_t {
_1_0,
};
/* This `BackwardLabel` struct used for retaining a location for PC-Relative instructions.
/* This `BackwardLabel` struct is used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `below` an instruction that uses it.
* Which means that a branch would jump backwards.
*/
@@ -520,13 +515,11 @@ struct BackwardLabel {
uint8_t* Location {};
};
/* This `SingleUseForwardLabel` struct used for retaining a location for PC-Relative instructions.
/* This `ForwardLabel` struct is used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `above` an instruction that uses it.
* Which means that a branch would jump forwards.
*
* The `ForwardLabel` struct can be bound to multiple instructions, so it needs a vector for each bind instruction type.
*/
struct SingleUseForwardLabel {
struct ForwardLabel {
enum class InstType {
UNKNOWN,
ADR,
@@ -537,12 +530,16 @@ struct SingleUseForwardLabel {
RELATIVE_LOAD,
LONG_ADDRESS_GEN,
};
uint8_t* Location {};
InstType Type = InstType::UNKNOWN;
};
struct ForwardLabel {
fextl::vector<SingleUseForwardLabel> Insts {};
struct Reference {
uint8_t* Location {};
InstType Type = InstType::UNKNOWN;
};
// The first element is stored separately to avoid allocations for simple cases
Reference FirstInst;
fextl::vector<Reference> Insts;
};
/* This `BiDirectionalLabel` struct used for retaining a location for PC-Relative instructions.
@@ -554,14 +551,12 @@ struct BiDirectionalLabel {
ForwardLabel Forward;
};
static inline void AddLocationToLabel(SingleUseForwardLabel* Label, SingleUseForwardLabel&& Location) {
LOGMAN_THROW_A_FMT(Label->Type == SingleUseForwardLabel::InstType::UNKNOWN, "Trying to bind a SingleUseForwardLabel to multiple "
"locations. Use ForwardLabel instead.");
*Label = std::move(Location);
}
static inline void AddLocationToLabel(ForwardLabel* Label, SingleUseForwardLabel&& Location) {
Label->Insts.emplace_back(std::move(Location));
static inline void AddLocationToLabel(ForwardLabel* Label, ForwardLabel::Reference&& Location) {
if (Label->FirstInst.Location == nullptr) {
Label->FirstInst = Location;
} else {
Label->Insts.push_back(Location);
}
}
// Some FCMA ASIMD instructions support a rotation argument.
@@ -630,15 +625,15 @@ public:
// Bind a backward label to an address.
// Address that is bound is the current emitter location.
void Bind(BackwardLabel* Label) {
LOGMAN_THROW_AA_FMT(Label->Location == nullptr, "Trying to bind a label twice");
LOGMAN_THROW_A_FMT(Label->Location == nullptr, "Trying to bind a label twice");
Label->Location = GetCursorAddress<uint8_t*>();
}
void Bind(const SingleUseForwardLabel* Label) {
void Bind(const ForwardLabel::Reference* Label) {
uint8_t* CurrentAddress = GetCursorAddress<uint8_t*>();
// Patch up the instructions
switch (Label->Type) {
case SingleUseForwardLabel::InstType::ADR: {
case ForwardLabel::InstType::ADR: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
@@ -650,7 +645,7 @@ public:
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::ADRP: {
case ForwardLabel::InstType::ADRP: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
@@ -664,7 +659,7 @@ public:
break;
}
case SingleUseForwardLabel::InstType::B: {
case ForwardLabel::InstType::B: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -678,7 +673,7 @@ public:
break;
}
case SingleUseForwardLabel::InstType::TEST_BRANCH: {
case ForwardLabel::InstType::TEST_BRANCH: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -691,8 +686,8 @@ public:
break;
}
case SingleUseForwardLabel::InstType::BC:
case SingleUseForwardLabel::InstType::RELATIVE_LOAD: {
case ForwardLabel::InstType::BC:
case ForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -704,7 +699,7 @@ public:
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::LONG_ADDRESS_GEN: {
case ForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
@@ -749,10 +744,9 @@ public:
// Bind a forward label to a location.
// This walks all the instructions in the label's vector.
// Then backpatching all instructions that have used the label.
template<bool WarnAboutEmpty = false>
void Bind(ForwardLabel* Label) {
if constexpr (WarnAboutEmpty) {
LOGMAN_THROW_A_FMT(Label->Insts.empty() == false, "Binding forward label that didn't have any instructions using it");
if (Label->FirstInst.Location) {
Bind(&Label->FirstInst);
}
for (auto& Inst : Label->Insts) {
Bind(&Inst);
@@ -765,12 +759,18 @@ public:
if (!Label->Backward.Location) {
Bind(&Label->Backward);
}
Bind<false>(&Label->Forward);
Bind(&Label->Forward);
}
#include <CodeEmitter/VixlUtils.inl>
public:
// This symbol is used to allow external tooling (IDEs, clang-format, ...) to process the included files individually:
// If defined, the files will inject member functions into this class.
// If not, the files will wrap the member functions in a class so that tooling will process them properly.
#define INCLUDED_BY_EMITTER
// TODO: Implement SME when it matters.
#include <CodeEmitter/ALUOps.inl>
#include <CodeEmitter/BranchOps.inl>
@@ -780,7 +780,9 @@ public:
#include <CodeEmitter/ASIMDOps.inl>
#include <CodeEmitter/SVEOps.inl>
private:
#undef INCLUDED_BY_EMITTER
protected:
template<typename T>
uint32_t Encode_ra(T Reg) const {
return Reg.Idx() << 10;
@@ -792,7 +794,6 @@ private:
uint32_t Encode_rt2(T Reg) const {
return Reg.Idx() << 10;
}
template<>
uint32_t Encode_rt2(uint32_t Reg) const {
return Reg << 10;
}
@@ -828,7 +829,6 @@ private:
uint32_t Encode_rt(T Reg) const {
return Reg.Idx();
}
template<>
uint32_t Encode_rt(Prefetch Reg) const {
return FEXCore::ToUnderlying(Reg);
}
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
+160 -210
View File
@@ -16,17 +16,25 @@
* Exceptions to this rule will have asserts in the emitter implementation when misused.
*
*/
#pragma once
#ifndef INCLUDED_BY_EMITTER
#include <CodeEmitter/Emitter.h>
namespace ARMEmitter {
struct EmitterOps : Emitter {
#endif
public:
// Advanced SIMD scalar copy
// Advanced SIMD scalar copy
void dup(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Index) {
constexpr uint32_t Op = 0b0101'1110'0000'0000'0000'01 << 10;
const uint32_t SizeImm = FEXCore::ToUnderlying(size);
const uint32_t IndexShift = SizeImm + 1;
const uint32_t ElementSize = 1U << SizeImm;
const uint32_t MaxIndex = 128U / (ElementSize * 8);
[[maybe_unused]] const uint32_t MaxIndex = 128U / (ElementSize * 8);
LOGMAN_THROW_AA_FMT(Index < MaxIndex, "Index too large. Index={}, Max Index: {}", Index, MaxIndex);
LOGMAN_THROW_A_FMT(Index < MaxIndex, "Index too large. Index={}, Max Index: {}", Index, MaxIndex);
const uint32_t imm5 = (Index << IndexShift) | ElementSize;
@@ -37,7 +45,7 @@ public:
dup(size, rd, rn, Index);
}
// Advanced SIMD scalar three same FP16
// Advanced SIMD scalar three same FP16
void fmulx(HRegister rd, HRegister rn, HRegister rm) {
ASIMDScalarThreeSameFP16(0, 0, 0b011, rm, rn, rd);
}
@@ -66,7 +74,7 @@ public:
ASIMDScalarThreeSameFP16(1, 1, 0b101, rm, rn, rd);
}
// Advanced SIMD scalar two-register miscellaneous FP16
// Advanced SIMD scalar two-register miscellaneous FP16
void fcvtns(HRegister rd, HRegister rn) {
ASIMDScalarTwoRegMiscFP16(0, 0, 0b11010, rn, rd);
}
@@ -128,9 +136,9 @@ public:
ASIMDScalarTwoRegMiscFP16(1, 1, 0b11101, rn, rd);
}
// Advanced SIMD scalar three same extra
// XXX:
// Advanced SIMD scalar two-register miscellaneous
// Advanced SIMD scalar three same extra
// XXX:
// Advanced SIMD scalar two-register miscellaneous
void suqadd(ScalarRegSize size, VRegister rd, VRegister rn) {
ASIMDScalar2RegMisc(0, 0, size, 0b00011, rd, rn);
}
@@ -140,67 +148,55 @@ public:
///< Comparison against 0.0
void cmgt(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMDScalar2RegMisc(0, 0, size, 0b01000, rd, rn);
}
///< Comparison against 0.0
void cmeq(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMDScalar2RegMisc(0, 0, size, 0b01001, rd, rn);
}
///< Comparison against 0.0
void cmlt(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMDScalar2RegMisc(0, 0, size, 0b01010, rd, rn);
}
void abs(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMDScalar2RegMisc(0, 0, size, 0b01011, rd, rn);
}
///< size is destination size.
void sqxtn(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size != ScalarRegSize::i64Bit, "64-bit destination not supported");
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "64-bit destination not supported");
ASIMDScalar2RegMisc(0, 0, size, 0b10100, rd, rn);
}
void fcvtns(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 0, ConvertedSize, 0b11010, rd, rn);
}
void fcvtms(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 0, ConvertedSize, 0b11011, rd, rn);
}
void fcvtas(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 0, ConvertedSize, 0b11100, rd, rn);
}
void scvtf(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 0, ConvertedSize, 0b11101, rd, rn);
}
@@ -249,70 +245,58 @@ public:
}
///< Comparison against 0.0
void cmge(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMDScalar2RegMisc(0, 1, size, 0b01000, rd, rn);
}
///< Comparison against 0.0
void cmle(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMDScalar2RegMisc(0, 1, size, 0b01001, rd, rn);
}
void neg(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMDScalar2RegMisc(0, 1, size, 0b01011, rd, rn);
}
///< size is destination.
void sqxtun(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size != ScalarRegSize::i64Bit, "64-bit destination not supported");
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "64-bit destination not supported");
ASIMDScalar2RegMisc(0, 1, size, 0b10010, rd, rn);
}
///< size is destination.
void uqxtn(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size != ScalarRegSize::i64Bit, "64-bit destination not supported");
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "64-bit destination not supported");
ASIMDScalar2RegMisc(0, 1, size, 0b10100, rd, rn);
}
///< size is destination.
void fcvtxn(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
ASIMDScalar2RegMisc(0, 1, ScalarRegSize::i16Bit, 0b10110, rd, rn);
}
void fcvtnu(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 1, ConvertedSize, 0b11010, rd, rn);
}
void fcvtmu(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 1, ConvertedSize, 0b11011, rd, rn);
}
void fcvtau(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 1, ConvertedSize, 0b11100, rd, rn);
}
void ucvtf(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(0, 1, ConvertedSize, 0b11101, rd, rn);
}
@@ -366,73 +350,55 @@ public:
}
void fmaxnmp(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(1, 1, ConvertedSize, 0b01100, rd, rn);
}
void faddp(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(1, 1, ConvertedSize, 0b01101, rd, rn);
}
void fmaxp(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMDScalar2RegMisc(1, 1, ConvertedSize, 0b01111, rd, rn);
}
void fminnmp(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
ASIMDScalar2RegMisc(1, 1, size, 0b01100, rd, rn);
}
void fminp(ScalarRegSize size, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
ASIMDScalar2RegMisc(1, 1, size, 0b01111, rd, rn);
}
// Advanced SIMD scalar three different
// Advanced SIMD scalar three different
///< size is destination.
void sqdmlal(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i32Bit :
ScalarRegSize::i16Bit;
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i32Bit : ScalarRegSize::i16Bit;
ASIMD3RegDifferent(0, ConvertedSize, 0b1001, rd, rn, rm);
}
///< size is destination.
void sqdmlsl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i32Bit :
ScalarRegSize::i16Bit;
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i32Bit : ScalarRegSize::i16Bit;
ASIMD3RegDifferent(0, ConvertedSize, 0b1011, rd, rn, rm);
}
///< size is destination.
void sqdmull(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i32Bit :
ScalarRegSize::i16Bit;
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i32Bit : ScalarRegSize::i16Bit;
ASIMD3RegDifferent(0, ConvertedSize, 0b1101, rd, rn, rm);
}
// Advanced SIMD scalar three same
// Advanced SIMD scalar three same
void sqadd(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
ASIMD3RegSame(0, size, 0b00001, rd, rn, rm);
}
@@ -440,71 +406,62 @@ public:
ASIMD3RegSame(0, size, 0b00101, rd, rn, rm);
}
void cmgt(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(0, size, 0b00110, rd, rn, rm);
}
void cmge(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(0, size, 0b00111, rd, rn, rm);
}
void sshl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(0, size, 0b01000, rd, rn, rm);
}
void sqshl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
ASIMD3RegSame(0, size, 0b01001, rd, rn, rm);
}
void srshl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(0, size, 0b01010, rd, rn, rm);
}
void sqrshl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
ASIMD3RegSame(0, size, 0b01011, rd, rn, rm);
}
void add(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(0, size, 0b10000, rd, rn, rm);
}
void cmtst(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(0, size, 0b10001, rd, rn, rm);
}
void sqdmulh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i32Bit || size == ScalarRegSize::i16Bit, "Invalid size");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i32Bit || size == ScalarRegSize::i16Bit, "Invalid size");
ASIMD3RegSame(0, size, 0b10110, rd, rn, rm);
}
void fmulx(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMD3RegSame(0, ConvertedSize, 0b11011, rd, rn, rm);
}
void fcmeq(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMD3RegSame(0, ConvertedSize, 0b11100, rd, rn, rm);
}
void frecps(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMD3RegSame(0, ConvertedSize, 0b11111, rd, rn, rm);
}
void frsqrts(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
ASIMD3RegSame(0, size, 0b11111, rd, rn, rm);
}
void uqadd(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
@@ -514,75 +471,69 @@ public:
ASIMD3RegSame(1, size, 0b00101, rd, rn, rm);
}
void cmhi(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(1, size, 0b00110, rd, rn, rm);
}
void cmhs(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(1, size, 0b00111, rd, rn, rm);
}
void ushl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(1, size, 0b01000, rd, rn, rm);
}
void uqshl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
ASIMD3RegSame(1, size, 0b01001, rd, rn, rm);
}
void urshl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(1, size, 0b01010, rd, rn, rm);
}
void uqrshl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
ASIMD3RegSame(1, size, 0b01011, rd, rn, rm);
}
void sub(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(1, size, 0b10000, rd, rn, rm);
}
void cmeq(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit, "Only supports 64-bit");
ASIMD3RegSame(1, size, 0b10001, rd, rn, rm);
}
void sqrdmulh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i32Bit || size == ScalarRegSize::i16Bit, "Invalid size");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i32Bit || size == ScalarRegSize::i16Bit, "Invalid size");
ASIMD3RegSame(1, size, 0b10110, rd, rn, rm);
}
void fcmge(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMD3RegSame(1, ConvertedSize, 0b11100, rd, rn, rm);
}
void facge(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
const ScalarRegSize ConvertedSize =
size == ScalarRegSize::i64Bit ?
ScalarRegSize::i16Bit :
ScalarRegSize::i8Bit;
const ScalarRegSize ConvertedSize = size == ScalarRegSize::i64Bit ? ScalarRegSize::i16Bit : ScalarRegSize::i8Bit;
ASIMD3RegSame(1, ConvertedSize, 0b11101, rd, rn, rm);
}
void fabd(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
ASIMD3RegSame(1, size, 0b11010, rd, rn, rm);
}
void fcmgt(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
ASIMD3RegSame(1, size, 0b11100, rd, rn, rm);
}
void facgt(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for float convert");
ASIMD3RegSame(1, size, 0b11101, rd, rn, rm);
}
// Advanced SIMD scalar shift by immediate
// Advanced SIMD scalar shift by immediate
void sshr(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -592,8 +543,8 @@ public:
ASIMDScalarShiftByImm(0, immh, immb, 0b00000, rd, rn);
}
void ssra(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -603,8 +554,8 @@ public:
ASIMDScalarShiftByImm(0, immh, immb, 0b00010, rd, rn);
}
void srshr(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -614,8 +565,8 @@ public:
ASIMDScalarShiftByImm(0, immh, immb, 0b00100, rd, rn);
}
void srsra(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -625,8 +576,8 @@ public:
ASIMDScalarShiftByImm(0, immh, immb, 0b00110, rd, rn);
}
void shl(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
// Shift encoded a bit weirdly.
// shift = immh:immb - elementsize but immh is /also/ used for element size.
const uint32_t immh = 1 << FEXCore::ToUnderlying(size) | (Shift >> 3);
@@ -644,7 +595,7 @@ public:
///< size is destination
void sqshrn(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < ScalarRegSizeInBits(size), "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqshrn");
LOGMAN_THROW_A_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqshrn");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -655,7 +606,7 @@ public:
}
void sqrshrn(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < ScalarRegSizeInBits(size), "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqshrn");
LOGMAN_THROW_A_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqshrn");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -666,8 +617,8 @@ public:
}
// TODO: SCVTF, FCVTZS
void ushr(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -677,8 +628,8 @@ public:
ASIMDScalarShiftByImm(1, immh, immb, 0b00000, rd, rn);
}
void usra(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -688,8 +639,8 @@ public:
ASIMDScalarShiftByImm(1, immh, immb, 0b00010, rd, rn);
}
void urshr(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -699,8 +650,8 @@ public:
ASIMDScalarShiftByImm(1, immh, immb, 0b00100, rd, rn);
}
void ursra(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -710,8 +661,8 @@ public:
ASIMDScalarShiftByImm(1, immh, immb, 0b00110, rd, rn);
}
void sri(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -721,8 +672,8 @@ public:
ASIMDScalarShiftByImm(1, immh, immb, 0b01000, rd, rn);
}
void sli(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_AA_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < 64, "Invalid shift for sshr");
LOGMAN_THROW_A_FMT(size == ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sshr");
// Shift encoded a bit weirdly.
// shift = immh:immb - elementsize but immh is /also/ used for element size.
const uint32_t immh = 1 << FEXCore::ToUnderlying(size) | (Shift >> 3);
@@ -748,7 +699,7 @@ public:
///< size is destination.
void sqshrun(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < ScalarRegSizeInBits(size), "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqshrun");
LOGMAN_THROW_A_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqshrun");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -760,7 +711,7 @@ public:
///< size is destination.
void sqrshrun(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < ScalarRegSizeInBits(size), "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqrshrun");
LOGMAN_THROW_A_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqrshrun");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -772,7 +723,7 @@ public:
///< size is destination.
void uqshrn(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < ScalarRegSizeInBits(size), "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqrshrun");
LOGMAN_THROW_A_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqrshrun");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -784,7 +735,7 @@ public:
///< size is destination.
void uqrshrn(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Shift) {
LOGMAN_THROW_A_FMT(Shift > 0 && Shift < ScalarRegSizeInBits(size), "Invalid shift for sshr");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqrshrun");
LOGMAN_THROW_A_FMT(size != ARMEmitter::ScalarRegSize::i64Bit, "Invalid size selected for sqrshrun");
const size_t SubregSizeInBits = ScalarRegSizeInBits(size);
// Shift encoded in immh:immb, but inverted with 128-bit source
// shift = (esize * 2) - immh:immb
@@ -794,10 +745,10 @@ public:
ASIMDScalarShiftByImm(1, immh, immb, 0b10011, rd, rn);
}
// TODO: UCVTF, FCVTZU
// Advanced SIMD scalar x indexed element
// XXX:
//
// Floating-point data-processing (1 source)
// Advanced SIMD scalar x indexed element
// XXX:
//
// Floating-point data-processing (1 source)
void fmov(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b000000, rd, rn);
}
@@ -991,14 +942,14 @@ public:
Float1Source(0, 0, 0b11, 0b001111, rd.V(), rn.V());
}
// Floating-point compare
// Floating-point compare
void fcmp(ScalarRegSize Size, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(Size != ScalarRegSize::i8Bit, "8-bit destination not supported");
LOGMAN_THROW_A_FMT(Size != ScalarRegSize::i8Bit, "8-bit destination not supported");
const auto ConvertedSize =
Size == ARMEmitter::ScalarRegSize::i64Bit ? 0b01 :
Size == ARMEmitter::ScalarRegSize::i32Bit ? 0b00 :
Size == ARMEmitter::ScalarRegSize::i16Bit ? 0b11 : 0;
const auto ConvertedSize = Size == ARMEmitter::ScalarRegSize::i64Bit ? 0b01 :
Size == ARMEmitter::ScalarRegSize::i32Bit ? 0b00 :
Size == ARMEmitter::ScalarRegSize::i16Bit ? 0b11 :
0;
FloatCompare(0, 0, ConvertedSize, 0b00, 0b00000, rn, rm);
}
@@ -1051,7 +1002,7 @@ public:
FloatCompare(0, 0, 0b11, 0b00, 0b11000, rn.V(), VReg::v0);
}
// Floating-point immediate
// Floating-point immediate
void fmov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, float Value) {
uint32_t M = 0;
uint32_t S = 0;
@@ -1061,16 +1012,13 @@ public:
if (size == ARMEmitter::ScalarRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
FEX_UNREACHABLE;
}
else if (size == ARMEmitter::ScalarRegSize::i32Bit) {
} else if (size == ARMEmitter::ScalarRegSize::i32Bit) {
ptype = 0b00;
imm8 = FP32ToImm8(Value);
}
else if (size == ARMEmitter::ScalarRegSize::i64Bit) {
} else if (size == ARMEmitter::ScalarRegSize::i64Bit) {
ptype = 0b01;
imm8 = FP64ToImm8(Value);
}
else {
} else {
FEX_UNREACHABLE;
}
@@ -1090,7 +1038,7 @@ public:
dc32(Instr);
}
// Floating-point conditional compare
// Floating-point conditional compare
void fccmp(SRegister rn, SRegister rm, StatusFlags flags, Condition Cond) {
FloatConditionalCompare(0, 0, 0b00, 0b0, rn.V(), rm.V(), flags, Cond);
}
@@ -1110,7 +1058,7 @@ public:
FloatConditionalCompare(0, 0, 0b11, 0b1, rn.V(), rm.V(), flags, Cond);
}
// Floating-point data-processing (2 source)
// Floating-point data-processing (2 source)
void fmul(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
Float2Source(size, 0, 0, 0b0000, rd, rn, rm);
}
@@ -1225,11 +1173,10 @@ public:
// Floating-point conditional select
void fcsel(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, Condition Cond) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for {}", __func__);
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit,
"Invalid size selected for {}", __func__);
const uint32_t ConvertedSize =
size == ScalarRegSize::i64Bit ? 0b01 :
size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
const uint32_t ConvertedSize = size == ScalarRegSize::i64Bit ? 0b01 : size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
FloatConditionalSelect(0, 0, ConvertedSize, rd, rn, rm, Cond);
}
@@ -1244,7 +1191,7 @@ public:
FloatConditionalSelect(0, 0, 0b11, rd.V(), rn.V(), rm.V(), Cond);
}
// Floating-point data-processing (3 source)
// Floating-point data-processing (3 source)
void fmadd(SRegister rd, SRegister rn, SRegister rm, SRegister ra) {
Float3Source(0, 0, 0b00, 0, 0, rd.V(), rn.V(), rm.V(), ra.V());
}
@@ -1285,7 +1232,7 @@ public:
}
private:
// Advanced SIMD scalar copy
// Advanced SIMD scalar copy
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn) {
uint32_t Instr = Op;
@@ -1297,7 +1244,7 @@ private:
dc32(Instr);
}
// Advanced SIMD scalar three same FP16
// Advanced SIMD scalar three same FP16
void ASIMDScalarThreeSameFP16(uint32_t U, uint32_t a, uint32_t opcode, HRegister rm, HRegister rn, HRegister rd) {
uint32_t Instr = 0b0101'1110'0100'0000'0000'0100'0000'0000;
@@ -1309,7 +1256,7 @@ private:
Instr |= rd.Idx();
dc32(Instr);
}
// Advanced SIMD scalar two-register miscellaneous FP16
// Advanced SIMD scalar two-register miscellaneous FP16
void ASIMDScalarTwoRegMiscFP16(uint32_t U, uint32_t a, uint32_t opcode, HRegister rn, HRegister rd) {
uint32_t Instr = 0b0101'1110'0111'1000'0000'1000'0000'0000;
@@ -1321,9 +1268,9 @@ private:
dc32(Instr);
}
// Advanced SIMD scalar three same extra
// XXX:
// Advanced SIMD scalar two-register miscellaneous
// Advanced SIMD scalar three same extra
// XXX:
// Advanced SIMD scalar two-register miscellaneous
void ASIMDScalar2RegMisc(uint32_t b20, uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0101'1110'0010'0000'0000'1000'0000'0000;
@@ -1336,9 +1283,9 @@ private:
dc32(Instr);
}
// Advanced SIMD scalar pairwise
// XXX:
// Advanced SIMD scalar three different
// Advanced SIMD scalar pairwise
// XXX:
// Advanced SIMD scalar three different
void ASIMD3RegDifferent(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0101'1110'0010'0000'0000'0000'0000'0000;
@@ -1350,7 +1297,7 @@ private:
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Advanced SIMD scalar three same
// Advanced SIMD scalar three same
void ASIMD3RegSame(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0101'1110'0010'0000'0000'0100'0000'0000;
@@ -1362,7 +1309,7 @@ private:
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Advanced SIMD scalar shift by immediate
// Advanced SIMD scalar shift by immediate
void ASIMDScalarShiftByImm(uint32_t U, uint32_t immh, uint32_t immb, uint32_t opcode, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0101'1111'0000'0000'0000'0100'0000'0000;
@@ -1374,9 +1321,9 @@ private:
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Advanced SIMD scalar x indexed element
// XXX:
// Floating-point data-processing (1 source)
// Advanced SIMD scalar x indexed element
// XXX:
// Floating-point data-processing (1 source)
void Float1Source(uint32_t M, uint32_t S, uint32_t ptype, uint32_t opcode, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0001'1110'0010'0000'0100'0000'0000'0000;
@@ -1390,16 +1337,15 @@ private:
dc32(Instr);
}
void Float1Source(ScalarRegSize size, uint32_t M, uint32_t S, uint32_t opcode, VRegister rd, VRegister rn) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for {}", __func__);
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit,
"Invalid size selected for {}", __func__);
const uint32_t ConvertedSize =
size == ScalarRegSize::i64Bit ? 0b01 :
size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
const uint32_t ConvertedSize = size == ScalarRegSize::i64Bit ? 0b01 : size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
Float1Source(M, S, ConvertedSize, opcode, rd, rn);
}
// Floating-point compare
// Floating-point compare
void FloatCompare(uint32_t M, uint32_t S, uint32_t ftype, uint32_t op, uint32_t opcode2, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0001'1110'0010'0000'0010'0000'0000'0000;
@@ -1413,9 +1359,9 @@ private:
dc32(Instr);
}
// Floating-point immediate
// XXX:
// Floating-point conditional compare
// Floating-point immediate
// XXX:
// Floating-point conditional compare
void FloatConditionalCompare(uint32_t M, uint32_t S, uint32_t ptype, uint32_t op, VRegister rn, VRegister rm, StatusFlags flags, Condition Cond) {
uint32_t Instr = 0b0001'1110'0010'0000'0000'0100'0000'0000;
@@ -1430,7 +1376,7 @@ private:
dc32(Instr);
}
// Floating-point data-processing (2 source)
// Floating-point data-processing (2 source)
void Float2Source(uint32_t M, uint32_t S, uint32_t ptype, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0001'1110'0010'0000'0000'1000'0000'0000;
@@ -1447,16 +1393,15 @@ private:
}
void Float2Source(ScalarRegSize size, uint32_t M, uint32_t S, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_AA_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit, "Invalid size selected for {}", __func__);
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i64Bit || size == ScalarRegSize::i32Bit,
"Invalid size selected for {}", __func__);
const uint32_t ConvertedSize =
size == ScalarRegSize::i64Bit ? 0b01 :
size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
const uint32_t ConvertedSize = size == ScalarRegSize::i64Bit ? 0b01 : size == ScalarRegSize::i32Bit ? 0b00 : 0b11;
Float2Source(M, S, ConvertedSize, opcode, rd, rn, rm);
}
// Floating-point conditional select
// Floating-point conditional select
void FloatConditionalSelect(uint32_t M, uint32_t S, uint32_t ptype, VRegister rd, VRegister rn, VRegister rm, Condition Cond) {
uint32_t Instr = 0b0001'1110'0010'0000'0000'1100'0000'0000;
@@ -1470,7 +1415,7 @@ private:
dc32(Instr);
}
// Floating-point data-processing (3 source)
// Floating-point data-processing (3 source)
void Float3Source(uint32_t M, uint32_t S, uint32_t ptype, uint32_t o1, uint32_t o0, VRegister rd, VRegister rn, VRegister rm, VRegister ra) {
uint32_t Instr = 0b0001'1111'0000'0000'0000'0000'0000'0000;
@@ -1485,3 +1430,8 @@ private:
Instr |= Encode_rd(rd);
dc32(Instr);
}
#ifndef INCLUDED_BY_EMITTER
}; // struct LoadstoreEmitterOps
} // namespace ARMEmitter
#endif
+160 -148
View File
@@ -4,173 +4,185 @@
* This is mostly a mashup of various instruction types.
* Nothing follows an explicit pattern since they are mostly different.
*/
#pragma once
#ifndef INCLUDED_BY_EMITTER
#include <CodeEmitter/Emitter.h>
namespace ARMEmitter {
struct EmitterOps : Emitter {
#endif
public:
// System with result
// TODO: SYSL
// System Instruction
// TODO: AT
// TODO: CFP
// TODO: CPP
void dc(ARMEmitter::DataCacheOperation DCOp, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0000'1000'0111 << 12;
SystemInstruction(Op, 0, FEXCore::ToUnderlying(DCOp), rt);
}
// TODO: DVP
// TODO: IC
// TODO: TLBI
// System with result
// TODO: SYSL
// System Instruction
// TODO: AT
// TODO: CFP
// TODO: CPP
void dc(ARMEmitter::DataCacheOperation DCOp, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0000'1000'0111 << 12;
SystemInstruction(Op, 0, FEXCore::ToUnderlying(DCOp), rt);
}
// TODO: DVP
// TODO: IC
// TODO: TLBI
// Exception generation
void svc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b01, Imm);
}
void hvc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b10, Imm);
}
void smc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b11, Imm);
}
void brk(uint32_t Imm) {
ExceptionGeneration(0b001, 0b000, 0b00, Imm);
}
void hlt(uint32_t Imm) {
ExceptionGeneration(0b010, 0b000, 0b00, Imm);
}
void tcancel(uint32_t Imm) {
ExceptionGeneration(0b011, 0b000, 0b00, Imm);
}
void dcps1(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b01, Imm);
}
void dcps2(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b10, Imm);
}
void dcps3(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b11, Imm);
}
// System instructions with register argument
void wfet(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b000, rt);
}
void wfit(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b001, rt);
}
// Exception generation
void svc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b01, Imm);
}
void hvc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b10, Imm);
}
void smc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b11, Imm);
}
void brk(uint32_t Imm) {
ExceptionGeneration(0b001, 0b000, 0b00, Imm);
}
void hlt(uint32_t Imm) {
ExceptionGeneration(0b010, 0b000, 0b00, Imm);
}
void tcancel(uint32_t Imm) {
ExceptionGeneration(0b011, 0b000, 0b00, Imm);
}
void dcps1(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b01, Imm);
}
void dcps2(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b10, Imm);
}
void dcps3(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b11, Imm);
}
// System instructions with register argument
void wfet(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b000, rt);
}
void wfit(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b001, rt);
}
// Hints
void nop() {
Hint(ARMEmitter::HintRegister::NOP);
}
void yield() {
Hint(ARMEmitter::HintRegister::YIELD);
}
void wfe() {
Hint(ARMEmitter::HintRegister::WFE);
}
void wfi() {
Hint(ARMEmitter::HintRegister::WFI);
}
void sev() {
Hint(ARMEmitter::HintRegister::SEV);
}
void sevl() {
Hint(ARMEmitter::HintRegister::SEVL);
}
void dgh() {
Hint(ARMEmitter::HintRegister::DGH);
}
void csdb() {
Hint(ARMEmitter::HintRegister::CSDB);
}
// Hints
void nop() {
Hint(ARMEmitter::HintRegister::NOP);
}
void yield() {
Hint(ARMEmitter::HintRegister::YIELD);
}
void wfe() {
Hint(ARMEmitter::HintRegister::WFE);
}
void wfi() {
Hint(ARMEmitter::HintRegister::WFI);
}
void sev() {
Hint(ARMEmitter::HintRegister::SEV);
}
void sevl() {
Hint(ARMEmitter::HintRegister::SEVL);
}
void dgh() {
Hint(ARMEmitter::HintRegister::DGH);
}
void csdb() {
Hint(ARMEmitter::HintRegister::CSDB);
}
// Barriers
void clrex(uint32_t imm = 15) {
LOGMAN_THROW_AA_FMT(imm < 16, "Immediate out of range");
Barrier(ARMEmitter::BarrierRegister::CLREX, imm);
}
void dsb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
}
void dmb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
}
void isb() {
Barrier(ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(ARMEmitter::BarrierScope::SY));
}
void sb() {
Barrier(ARMEmitter::BarrierRegister::SB, 0);
}
void tcommit() {
Barrier(ARMEmitter::BarrierRegister::TCOMMIT, 0);
}
// Barriers
void clrex(uint32_t imm = 15) {
LOGMAN_THROW_A_FMT(imm < 16, "Immediate out of range");
Barrier(ARMEmitter::BarrierRegister::CLREX, imm);
}
void dsb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
}
void dmb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
}
void isb() {
Barrier(ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(ARMEmitter::BarrierScope::SY));
}
void sb() {
Barrier(ARMEmitter::BarrierRegister::SB, 0);
}
void tcommit() {
Barrier(ARMEmitter::BarrierRegister::TCOMMIT, 0);
}
// System register move
void msr(ARMEmitter::SystemRegister reg, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0001 << 20;
SystemRegisterMove(Op, rt, reg);
}
// System register move
void msr(ARMEmitter::SystemRegister reg, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0001 << 20;
SystemRegisterMove(Op, rt, reg);
}
void mrs(ARMEmitter::Register rd, ARMEmitter::SystemRegister reg) {
constexpr uint32_t Op = 0b1101'0101'0011 << 20;
SystemRegisterMove(Op, rd, reg);
}
void mrs(ARMEmitter::Register rd, ARMEmitter::SystemRegister reg) {
constexpr uint32_t Op = 0b1101'0101'0011 << 20;
SystemRegisterMove(Op, rd, reg);
}
private:
// Exception Generation
void ExceptionGeneration(uint32_t opc, uint32_t op2, uint32_t LL, uint32_t Imm) {
LOGMAN_THROW_AA_FMT((Imm & 0xFFFF'0000) == 0, "Imm amount too large");
// Exception Generation
void ExceptionGeneration(uint32_t opc, uint32_t op2, uint32_t LL, uint32_t Imm) {
LOGMAN_THROW_A_FMT((Imm & 0xFFFF'0000) == 0, "Imm amount too large");
uint32_t Instr = 0b1101'0100 << 24;
uint32_t Instr = 0b1101'0100 << 24;
Instr |= opc << 21;
Instr |= Imm << 5;
Instr |= op2 << 2;
Instr |= LL;
Instr |= opc << 21;
Instr |= Imm << 5;
Instr |= op2 << 2;
Instr |= LL;
dc32(Instr);
}
dc32(Instr);
}
// System instructions with register argument
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, ARMEmitter::Register rt) {
uint32_t Instr = 0b1101'0101'0000'0011'0001 << 12;
// System instructions with register argument
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, ARMEmitter::Register rt) {
uint32_t Instr = 0b1101'0101'0000'0011'0001 << 12;
Instr |= CRm << 8;
Instr |= op2 << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
Instr |= CRm << 8;
Instr |= op2 << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
// Hints
void Hint(ARMEmitter::HintRegister Reg) {
uint32_t Instr = 0b1101'0101'0000'0011'0010'0000'0001'1111U;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// Barriers
void Barrier(ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
uint32_t Instr = 0b1101'0101'0000'0011'0011'0000'0001'1111U;
Instr |= CRm << 8;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// Hints
void Hint(ARMEmitter::HintRegister Reg) {
uint32_t Instr = 0b1101'0101'0000'0011'0010'0000'0001'1111U;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// Barriers
void Barrier(ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
uint32_t Instr = 0b1101'0101'0000'0011'0011'0000'0001'1111U;
Instr |= CRm << 8;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// System Instruction
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, ARMEmitter::Register rt) {
uint32_t Instr = Op;
// System Instruction
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, ARMEmitter::Register rt) {
uint32_t Instr = Op;
Instr |= L << 21;
Instr |= SubOp;
Instr |= Encode_rt(rt);
Instr |= L << 21;
Instr |= SubOp;
Instr |= Encode_rt(rt);
dc32(Instr);
}
dc32(Instr);
}
// System register move
void SystemRegisterMove(uint32_t Op, ARMEmitter::Register rt, ARMEmitter::SystemRegister reg) {
uint32_t Instr = Op;
// System register move
void SystemRegisterMove(uint32_t Op, ARMEmitter::Register rt, ARMEmitter::SystemRegister reg) {
uint32_t Instr = Op;
Instr |= FEXCore::ToUnderlying(reg);
Instr |= Encode_rt(rt);
Instr |= FEXCore::ToUnderlying(reg);
Instr |= Encode_rt(rt);
dc32(Instr);
}
dc32(Instr);
}
#ifndef INCLUDED_BY_EMITTER
}; // struct LoadstoreEmitterOps
} // namespace ARMEmitter
#endif
+17 -21
View File
@@ -34,11 +34,7 @@
// by the corresponding fields in the logical instruction.
// If it can not be encoded, the function returns false, and the values pointed
// to by n, imm_s and imm_r are undefined.
static bool IsImmLogical(uint64_t value,
unsigned width,
unsigned* n = nullptr,
unsigned* imm_s = nullptr,
unsigned* imm_r = nullptr) {
static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr, unsigned* imm_s = nullptr, unsigned* imm_r = nullptr) {
[[maybe_unused]] constexpr auto kBRegSize = 8;
[[maybe_unused]] constexpr auto kHRegSize = 16;
[[maybe_unused]] constexpr auto kSRegSize = 32;
@@ -47,8 +43,7 @@ static bool IsImmLogical(uint64_t value,
constexpr auto kWRegSize = 32;
constexpr auto kXRegSize = 64;
LOGMAN_THROW_A_FMT((width == kBRegSize) || (width == kHRegSize) ||
(width == kSRegSize) || (width == kDRegSize), "Unexpected imm size");
LOGMAN_THROW_A_FMT((width == kBRegSize) || (width == kHRegSize) || (width == kSRegSize) || (width == kDRegSize), "Unexpected imm size");
bool negate = false;
@@ -182,12 +177,7 @@ static bool IsImmLogical(uint64_t value,
// (1 + 2^d + 2^(2d) + ...), i.e. 0x0001000100010001 or similar. These can
// be derived using a table lookup on CLZ(d).
static const uint64_t multipliers[] = {
0x0000000000000001UL,
0x0000000100000001UL,
0x0001000100010001UL,
0x0101010101010101UL,
0x1111111111111111UL,
0x5555555555555555UL,
0x0000000000000001UL, 0x0000000100000001UL, 0x0001000100010001UL, 0x0101010101010101UL, 0x1111111111111111UL, 0x5555555555555555UL,
};
uint64_t multiplier = multipliers[CountLeadingZeros(d, kXRegSize) - 57];
uint64_t candidate = (b - a) * multiplier;
@@ -244,7 +234,9 @@ static bool IsImmLogical(uint64_t value,
}
static inline bool IsIntN(unsigned n, int64_t x) {
if (n == 64) return true;
if (n == 64) {
return true;
}
int64_t limit = INT64_C(1) << (n - 1);
return (-limit <= x) && (x < limit);
}
@@ -271,11 +263,15 @@ V(57) V(58) V(59) V(60) V(61) V(62) V(63)
// clang-format on
#define DECLARE_IS_INT_N(N) \
static inline bool IsInt##N(int64_t x) { return IsIntN(N, x); }
#define DECLARE_IS_INT_N(N) \
static inline bool IsInt##N(int64_t x) { \
return IsIntN(N, x); \
}
#define DECLARE_IS_UINT_N(N) \
static inline bool IsUint##N(int64_t x) { return IsUintN(N, x); }
#define DECLARE_IS_UINT_N(N) \
static inline bool IsUint##N(int64_t x) { \
return IsUintN(N, x); \
}
INT_1_TO_63_LIST(DECLARE_IS_INT_N)
INT_1_TO_63_LIST(DECLARE_IS_UINT_N)
@@ -285,14 +281,14 @@ INT_1_TO_63_LIST(DECLARE_IS_UINT_N)
private:
template <typename V>
template<typename V>
static inline bool IsPowerOf2(V value) {
return (value != 0) && ((value & (value - 1)) == 0);
}
// Some compilers dislike negating unsigned integers,
// so we provide an equivalent.
template <typename T>
template<typename T>
static inline T UnsignedNegate(T value) {
static_assert(std::is_unsigned<T>::value);
return ~value + 1;
@@ -302,7 +298,7 @@ static inline uint64_t LowestSetBit(uint64_t value) {
return value & UnsignedNegate(value);
}
template <typename V>
template<typename V>
static inline int CountLeadingZeros(V value, int width = (sizeof(V) * 8)) {
#if COMPILER_HAS_BUILTIN_CLZ
if (width == 32) {
+1 -136
View File
@@ -9,25 +9,6 @@
"@PREFIX_LIB@/libGL.so.1.7.0"
]
},
"GLESv2": {
"Library": "libGLESv2-guest.so",
"Depends": [
"X11"
],
"Overlay": [
"@PREFIX_LIB@/libGLESv2.so",
"@PREFIX_LIB@/libGLESv2.so.2",
"@PREFIX_LIB@/libGLESv2.so.2.0.0"
]
},
"X11": {
"Library": "libX11-guest.so",
"Overlay": [
"@PREFIX_LIB@/libX11.so",
"@PREFIX_LIB@/libX11.so.6",
"@PREFIX_LIB@/libX11.so.6.4.0"
]
},
"Vulkan": {
"Library": "libvulkan-guest.so",
"Overlay": [
@@ -36,89 +17,6 @@
"@HOME@/.local/share/Steam/ubuntu12_32/steam-runtime/pinned_libs_64/libvulkan.so.1"
]
},
"xcb": {
"Depends": [
"X11"
],
"Library": "libxcb-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb.so",
"@PREFIX_LIB@/libxcb.so.1",
"@PREFIX_LIB@/libxcb.so.1.1.0"
]
},
"xcb-dri2": {
"Library": "libxcb-dri2-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-dri2.so",
"@PREFIX_LIB@/libxcb-dri2.so.0",
"@PREFIX_LIB@/libxcb-dri2.so.0.0.0"
]
},
"xcb-dri3": {
"Library": "libxcb-dri3-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-dri3.so",
"@PREFIX_LIB@/libxcb-dri3.so.0",
"@PREFIX_LIB@/libxcb-dri3.so.0.0.0"
]
},
"xcb-xfixes": {
"Library": "libxcb-xfixes-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-xfixes.so",
"@PREFIX_LIB@/libxcb-xfixes.so.0",
"@PREFIX_LIB@/libxcb-xfixes.so.0.0.0"
]
},
"xcb-shm": {
"Library": "libxcb-shm-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-shm.so",
"@PREFIX_LIB@/libxcb-shm.so.0",
"@PREFIX_LIB@/libxcb-shm.so.0.0.0"
]
},
"xcb-sync": {
"Library": "libxcb-sync-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-sync.so",
"@PREFIX_LIB@/libxcb-sync.so.1",
"@PREFIX_LIB@/libxcb-sync.so.1.0.0"
]
},
"xcb-randr": {
"Library": "libxcb-randr-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-randr.so",
"@PREFIX_LIB@/libxcb-randr.so.0",
"@PREFIX_LIB@/libxcb-randr.so.0.1.0"
]
},
"xcb-present": {
"Library": "libxcb-present-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-present.so",
"@PREFIX_LIB@/libxcb-present.so.0",
"@PREFIX_LIB@/libxcb-present.so.0.0.0"
]
},
"xcb-glx": {
"Library": "libxcb-glx-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxcb-glx.so",
"@PREFIX_LIB@/libxcb-glx.so.0",
"@PREFIX_LIB@/libxcb-glx.so.0.0.0"
]
},
"xshmfence": {
"Library": "libxshmfence-guest.so",
"Overlay": [
"@PREFIX_LIB@/libxshmfence.so",
"@PREFIX_LIB@/libxshmfence.so.1",
"@PREFIX_LIB@/libxshmfence.so.1.0.0"
]
},
"drm": {
"Library": "libdrm-guest.so",
"Overlay": [
@@ -141,38 +39,6 @@
"@PREFIX_LIB@/libfex_thunk_test.so"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
"@PREFIX_LIB@/libXrender.so",
"@PREFIX_LIB@/libXrender.so.1",
"@PREFIX_LIB@/libXrender.so.1.3.0"
]
},
"Xext": {
"Library": "libXext-guest.so",
"Overlay": [
"@PREFIX_LIB@/libXext.so",
"@PREFIX_LIB@/libXext.so.6",
"@PREFIX_LIB@/libXext.so.6.4.0"
]
},
"Xfixes": {
"Library": "libXfixes-guest.so",
"Overlay": [
"@PREFIX_LIB@/libXfixes.so",
"@PREFIX_LIB@/libXfixes.so.3",
"@PREFIX_LIB@/libXfixes.so.3.1.0"
]
},
"OpenCL": {
"Library" : "libOpenCL-guest.so",
"Overlay": [
"@PREFIX_LIB@/libOpenCL.so",
"@PREFIX_LIB@/libOpenCL.so.1",
"@PREFIX_LIB@/libOpenCL.so.1.0.0"
]
},
"WaylandClient": {
"Library" : "libwayland-client-guest.so",
"Overlay": [
@@ -180,7 +46,6 @@
"@PREFIX_LIB@/libwayland-client.so.0",
"@PREFIX_LIB@/libwayland-client.so.0.20.0"
]
},
"":{}
}
}
}
+1 -1
+1 -1
View File
@@ -1,3 +1,3 @@
set(NAME tiny-json)
set(SRCS tiny-json.c)
add_library(${NAME} ${SRCS})
add_library(${NAME} STATIC ${SRCS})
Vendored Submodule
+1
Submodule External/tracy added at 5d542dc09f.
+1 -1
+1 -1
+15 -5
View File
@@ -188,27 +188,33 @@ def print_man_environment_tail():
# Additional environment variables that live outside of the normal loop
print_man_env_option(
"FEX_APP_CONFIG_LOCATION",
"APP_CONFIG_LOCATION",
[
"Allows the user to override where FEX looks for configuration files",
"By default FEX will look in {$HOME, $XDG_CONFIG_HOME}/.fex-emu/",
"This will override the full path",
"If FEX_PORTABLE is declared then relative paths are also supported",
"For FEXInterpreter: Relative to the FEXInterpreter binary",
"For WINE: Relative to %LOCALAPPDATA%"
],
"''", True)
print_man_env_option(
"FEX_APP_CONFIG",
"APP_CONFIG",
[
"Allows the user to override where FEX looks for only the application config file",
"By default FEX will look in {$HOME, $XDG_CONFIG_HOME}/.fex-emu/Config.json",
"This will override this file location",
"One must be careful with this option as it will override any applications that load with execve as well"
"If you need to support applications that execve then use FEX_APP_CONFIG_LOCATION instead"
"If FEX_PORTABLE is declared then relative paths are also supported",
"For FEXInterpreter: Relative to the FEXInterpreter binary",
"For WINE: Relative to %LOCALAPPDATA%"
],
"''", True)
print_man_env_option(
"FEX_APP_DATA_LOCATION",
"APP_DATA_LOCATION",
[
"Allows the user to override where FEX looks for data files",
"By default FEX will look in {$HOME, $XDG_DATA_HOME}/.fex-emu/",
@@ -218,9 +224,13 @@ def print_man_environment_tail():
"''", True)
print_man_env_option(
"FEX_PORTABLE",
"PORTABLE",
[
"Allows FEX to run without installation. Global locations for configuration and binfmt_misc are ignored. These files are instead read from <FEXInterpreterPath>/fex-emu/ by default.",
"Allows FEX to run without installation. Global locations for configuration and binfmt_misc are ignored.",
"For FEXInterpreter on Linux:",
"These files are instead read from <FEXInterpreterPath>/fex-emu/ by default.",
"For Arm64ec/Wow64 WINE builds:",
"These files are instead read from $LOCALAPPDATA/fex-emu/ by default.",
"For further customization, see FEX_APP_CONFIG_LOCATION and FEX_APP_DATA_LOCATION."
],
"''", True)
+19 -17
View File
@@ -44,7 +44,7 @@ class OpDefinition:
HasDest: bool
DestType: str
DestSize: str
NumElements: str
ElementSize: str
OpClass: str
HasSideEffects: bool
ImplicitFlagClobber: bool
@@ -67,7 +67,7 @@ class OpDefinition:
self.HasDest = False
self.DestType = None
self.DestSize = None
self.NumElements = None
self.ElementSize = None
self.OpClass = None
self.OpSize = 0
self.HasSideEffects = False
@@ -232,8 +232,8 @@ def parse_ops(ops):
if "DestSize" in op_val:
OpDef.DestSize = op_val["DestSize"]
if "NumElements" in op_val:
OpDef.NumElements = op_val["NumElements"]
if "ElementSize" in op_val:
OpDef.ElementSize = op_val["ElementSize"]
if len(op_class):
OpDef.OpClass = op_class
@@ -323,8 +323,8 @@ def print_ir_structs(defines):
output_file.write("struct __attribute__((packed)) IROp_Header {\n")
output_file.write("\tvoid* Data[0];\n")
output_file.write("\tIROps Op;\n\n")
output_file.write("\tuint8_t Size;\n")
output_file.write("\tuint8_t ElementSize;\n")
output_file.write("\tIR::OpSize Size;\n")
output_file.write("\tIR::OpSize ElementSize;\n")
output_file.write("\ttemplate<typename T>\n")
output_file.write("\tT const* C() const { return reinterpret_cast<T const*>(Data); }\n")
@@ -630,20 +630,19 @@ def print_ir_allocator_helpers():
output_file.write("\t\treturn IRPair<T>{Op, CreateNode(&Op->Header)};\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpSize(const OrderedNode *Op) const {\n")
output_file.write("\tIR::OpSize GetOpSize(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->Size;\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpElementSize(const OrderedNode *Op) const {\n")
output_file.write("\tIR::OpSize GetOpElementSize(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->ElementSize;\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpElements(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT(OpHasDest(Op), \"Op {} has no dest\\n\", GetName(HeaderOp->Op));\n")
output_file.write("\t\treturn HeaderOp->Size / HeaderOp->ElementSize;\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT(OpHasDest(Op), \"Op {} has no dest\\n\", GetOpName(Op));\n")
output_file.write("\t\treturn IR::OpSizeToSize(GetOpSize(Op)) / IR::OpSizeToSize(GetOpElementSize(Op));\n")
output_file.write("\t}\n\n")
output_file.write("\tbool OpHasDest(const OrderedNode *Op) const {\n")
@@ -699,8 +698,12 @@ def print_ir_allocator_helpers():
# We gather the "has x87?" flag as we go. This saves the user from
# having to keep track of whether they emitted any x87.
# Also changes the mmx state to X87.
if op.LoweredX87:
output_file.write("\t\tRecordX87Use();\n")
output_file.write(
"\t\tif(MMXState == MMXState_MMX) ChgStateMMX_X87();\n"
)
output_file.write("\t\tauto _Op = AllocateOp<IROp_{}, IROps::OP_{}>();\n".format(op.Name, op.Name.upper()))
@@ -724,11 +727,11 @@ def print_ir_allocator_helpers():
# We can only infer a size if we have arguments
if op.DestSize == None:
# We need to infer destination size
output_file.write("\t\tuint8_t InferSize = 0;\n")
output_file.write("\t\tIR::OpSize InferSize = OpSize::iUnsized;\n")
if len(op.Arguments) != 0:
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tuint8_t Size{} = GetOpSize({});\n".format(arg.Name, arg.Name))
output_file.write("\t\tauto Size{} = GetOpSize({});\n".format(arg.Name, arg.Name))
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tInferSize = std::max(InferSize, Size{});\n".format(arg.Name))
@@ -740,10 +743,10 @@ def print_ir_allocator_helpers():
if op.DestSize != None:
output_file.write("\t\t_Op.first->Header.Size = {};\n".format(op.DestSize))
if op.NumElements == None:
output_file.write("\t\t_Op.first->Header.ElementSize = _Op.first->Header.Size / ({});\n".format(1))
if op.ElementSize == None:
output_file.write("\t\t_Op.first->Header.ElementSize = _Op.first->Header.Size;\n")
else:
output_file.write("\t\t_Op.first->Header.ElementSize = _Op.first->Header.Size / ({});\n".format(op.NumElements))
output_file.write("\t\t_Op.first->Header.ElementSize = {};\n".format(op.ElementSize))
# Insert validation here
if op.EmitValidation != None:
@@ -826,4 +829,3 @@ print_ir_dispatcher_defs()
print_ir_dispatcher_dispatch()
output_dispatch_file.close()
+16 -12
View File
@@ -105,17 +105,18 @@ set (SRCS
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp
Interface/Core/JIT/Arm64/Arm64Relocations.cpp
Interface/Core/Interpreter/Fallbacks/StringCompareFallbacks.cpp
Interface/Core/JIT/JIT.cpp
Interface/Core/JIT/ALUOps.cpp
Interface/Core/JIT/AtomicOps.cpp
Interface/Core/JIT/BranchOps.cpp
Interface/Core/JIT/ConversionOps.cpp
Interface/Core/JIT/EncryptionOps.cpp
Interface/Core/JIT/MemoryOps.cpp
Interface/Core/JIT/MiscOps.cpp
Interface/Core/JIT/MoveOps.cpp
Interface/Core/JIT/VectorOps.cpp
Interface/Core/JIT/Arm64Relocations.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/H0F38Tables.cpp
@@ -126,7 +127,6 @@ set (SRCS
Interface/Core/X86Tables/SecondaryTables.cpp
Interface/Core/X86Tables/VEXTables.cpp
Interface/Core/X86Tables/X87Tables.cpp
Interface/Core/X86Tables/XOPTables.cpp
Interface/GDBJIT/GDBJIT.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
@@ -338,6 +338,10 @@ add_library(FEXCore_Base STATIC ${FEXCORE_BASE_SRCS})
target_link_libraries(FEXCore_Base ${LIBS})
AddDefaultOptionsToTarget(FEXCore_Base)
if (ENABLE_FEXCORE_PROFILER AND FEXCORE_PROFILER_BACKEND STREQUAL "TRACY")
target_link_libraries(FEXCore_Base TracyClient)
endif()
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
+3 -3
View File
@@ -21,12 +21,12 @@ struct BitSet final {
ElementType* Memory;
void Allocate(size_t Elements) {
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
}
void Realloc(size_t Elements) {
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
}
void Free() {
@@ -68,7 +68,7 @@ struct BitSetView final {
ElementType* Memory;
void GetView(BitSet<T>& Set, uint64_t ElementOffset) {
LOGMAN_THROW_AA_FMT((ElementOffset % MinimumSize) == 0, "Bitset view offset needs to be aligned to size of backing element");
LOGMAN_THROW_A_FMT((ElementOffset % MinimumSize) == 0, "Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
+26
View File
@@ -10,6 +10,8 @@
#include <cstring>
#include <stdint.h>
#include "Common/VectorRegType.h"
extern "C" {
#include "SoftFloat-3e/platform.h"
#include "SoftFloat-3e/softfloat.h"
@@ -233,6 +235,10 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
// Zero is a special case, the significand for +/- 0 is +/- zero.
if (lhs.Exponent == 0x0 && lhs.Significand == 0x0) {
return lhs;
}
X80SoftFloat Tmp = lhs;
Tmp.Exponent = 0x3FFF;
Tmp.Sign = lhs.Sign;
@@ -256,6 +262,12 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
// Zero is a special case, the exponent is always -inf
if (lhs.Exponent == 0x0 && lhs.Significand == 0x0) {
X80SoftFloat Result(1, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
int32_t TrueExp = lhs.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
#endif
@@ -466,6 +478,12 @@ struct FEX_PACKED X80SoftFloat {
return FEXCore::BitCast<double>(Result);
}
FEXCore::VectorRegType ToVector() const {
FEXCore::VectorRegType Ret {};
memcpy(&Ret, this, sizeof(*this));
return Ret;
}
LIBRARY_PRECISION ToFMax(softfloat_state* state) const {
#ifdef _WIN32
return ToF64(state);
@@ -557,12 +575,20 @@ struct FEX_PACKED X80SoftFloat {
*this = i32_to_extF80(rhs);
}
X80SoftFloat(const FEXCore::VectorRegType rhs) {
memcpy(this, &rhs, sizeof(*this));
}
void operator=(extFloat80_t rhs) {
Significand = rhs.signif;
Exponent = rhs.signExp & 0x7FFF;
Sign = rhs.signExp >> 15;
}
operator FEXCore::VectorRegType() const {
return ToVector();
}
operator extFloat80_t() const {
extFloat80_t Result {};
Result.signif = Significand;
+16
View File
@@ -0,0 +1,16 @@
// SPDX-License-Identifier: MIT
#pragma once
#ifdef _M_X86_64
#include <xmmintrin.h>
#endif
namespace FEXCore {
#ifdef _M_ARM_64
// Can't use uint8x16_t directly from arm_neon.h here.
// Overrides softfloat-3e's defines which causes problems.
using VectorRegType = __attribute__((neon_vector_type(16))) uint8_t;
#elif defined(_M_X86_64)
using VectorRegType = __m128i;
#endif
} // namespace FEXCore
+16 -6
View File
@@ -334,9 +334,14 @@ void ReloadMetaLayer() {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, ExpandedString);
} else if (!PathName->empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
fextl::string NamedRootFS = GetDataDirectory(false) + "RootFS/" + *PathName;
if (FHU::Filesystem::Exists(NamedRootFS)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, NamedRootFS);
const auto PathNameCopy = *PathName;
for (auto Global : {true, false}) {
for (auto DirectoryFetchers : {GetDataDirectory, GetConfigDirectory}) {
fextl::string NamedRootFS = DirectoryFetchers(Global) + "RootFS/" + PathNameCopy;
if (FHU::Filesystem::Exists(NamedRootFS)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, NamedRootFS);
}
}
}
}
}
@@ -356,9 +361,14 @@ void ReloadMetaLayer() {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, ExpandedString);
} else if (!PathName->empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
fextl::string NamedConfig = GetDataDirectory(false) + "ThunkConfigs/" + *PathName;
if (FHU::Filesystem::Exists(NamedConfig)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, NamedConfig);
const auto PathNameCopy = *PathName;
for (auto Global : {true, false}) {
for (auto DirectoryFetchers : {GetDataDirectory, GetConfigDirectory}) {
fextl::string NamedConfig = DirectoryFetchers(Global) + "ThunkConfigs/" + PathNameCopy;
if (FHU::Filesystem::Exists(NamedConfig)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, NamedConfig);
}
}
}
}
}
+16 -1
View File
@@ -3,7 +3,7 @@
"CPU": {
"Multiblock": {
"Type": "bool",
"Default": "false",
"Default": "true",
"ShortArg": "m",
"Desc": [
"Controls multiblock code compilation",
@@ -363,6 +363,14 @@
"Redirects the telemetry folder that FEX usually writes to.",
"By default telemetry data is stored in {$FEX_APP_DATA_LOCATION,{$XDG_DATA_HOME,$HOME}/.fex-emu/Telemetry/}"
]
},
"ProfileStats": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enables FEX's low-overhead sampling profile statistics.",
"Requires a supported version of Mangohud to see the results"
]
}
},
"Hacks": {
@@ -472,6 +480,13 @@
"Sleeps the process at startup for a duration of seconds.",
"Useful if an application crashes too quickly to attach a debugger."
]
},
"StartupSleepProcName": {
"Type": "str",
"Default": "",
"Desc": [
"Contrains the startup sleep to only apply to processes that match this name."
]
}
},
"Misc": {
@@ -24,14 +24,6 @@ fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNe
return fextl::make_unique<FEXCore::Context::ContextImpl>(Features);
}
void FEXCore::Context::ContextImpl::SetExitHandler(ExitHandler handler) {
CustomExitHandler = std::move(handler);
}
ExitHandler FEXCore::Context::ContextImpl::GetExitHandler() const {
return CustomExitHandler;
}
void FEXCore::Context::ContextImpl::CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
CompileBlock(Thread->CurrentFrame, GuestRIP);
}
+18 -39
View File
@@ -81,11 +81,6 @@ public:
// Context base class implementation.
bool InitCore() override;
void SetExitHandler(ExitHandler handler) override;
ExitHandler GetExitHandler() const override;
ExitReason RunUntilExit(FEXCore::Core::InternalThreadState* Thread) override;
void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) override;
void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) override;
@@ -93,8 +88,11 @@ public:
void HandleCallback(FEXCore::Core::InternalThreadState* Thread, uint64_t RIP) override;
bool IsAddressInCurrentBlock(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, uint64_t Size) override;
bool IsCurrentBlockSingleInst(FEXCore::Core::InternalThreadState* Thread) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) override;
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, uint64_t* HostGPRs, uint64_t PSTATE) override;
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, const uint64_t* HostGPRs, uint64_t PSTATE) override;
void SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, uint32_t EFLAGS) override;
void ReconstructXMMRegisters(const FEXCore::Core::InternalThreadState* Thread, __uint128_t* XMM_Low, __uint128_t* YMM_High) override;
@@ -113,33 +111,29 @@ public:
* Usecases:
* Parent thread Creation:
* - Thread = CreateThread(InitialRIP, InitialStack, nullptr, 0);
* - CTX->RunUntilExit(Thread);
* - CTX->ExecuteThread(Thread);
* OS thread Creation:
* - Thread = CreateThread(0, 0, NewState, PPID);
* - Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
* - ThreadHandler calls `CTX->ExecutionThread(Thread)`
* - ThreadHandler calls `CTX->ExecuteThread(Thread)`
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(0, 0, CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* - ExecuteThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(0, 0, NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState*
CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXCore::Core::CPUState* NewThreadState, uint64_t ParentTID) override;
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState* Thread) override;
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState* Thread, bool NeedsTLSUninstall) override;
void DestroyThread(FEXCore::Core::InternalThreadState* Thread) override;
#ifndef _WIN32
void LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) override;
@@ -235,8 +229,6 @@ public:
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
} Config;
std::atomic_bool CoreShuttingDown {false};
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
uint32_t StrictSplitLockMutex {};
@@ -249,8 +241,6 @@ public:
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
FEXCore::Context::ExitHandler CustomExitHandler;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
@@ -258,8 +248,6 @@ public:
~ContextImpl();
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestDestination,
FEXCore::Context::ExitFunctionLinkData* HostLink, const BlockDelinkerFunc& delinker);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, ExitFunctionLinkData* Record) {
@@ -281,7 +269,8 @@ public:
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
struct GenerateIRResult {
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
std::optional<IR::IRListView> IRView;
IR::RegisterAllocationData* RAData;
uint64_t TotalInstructions;
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
@@ -292,28 +281,17 @@ public:
struct CompileCodeResult {
void* CompiledCode;
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
FEXCore::Core::DebugData* DebugData;
bool GeneratedIR;
fextl::unique_ptr<FEXCore::Core::DebugData> DebugData;
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]]
CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
// Used for thread creation from syscalls
/**
* @brief Initializes TID, PID and TLS data for a thread
*
* @param Thread The internal FEX thread state object
*/
void InitializeThreadTLSData(FEXCore::Core::InternalThreadState* Thread);
void CopyMemoryMapping(FEXCore::Core::InternalThreadState* ParentThread, FEXCore::Core::InternalThreadState* ChildThread);
uint8_t GetGPRSize() const {
return Config.Is64BitMode ? 8 : 4;
IR::OpSize GetGPROpSize() const {
return Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
FEXCore::JITSymbols Symbols;
@@ -356,6 +334,10 @@ protected:
AtomicTSOEmulationEnabled = false;
VectorAtomicTSOEmulationEnabled = false;
MemcpyAtomicTSOEmulationEnabled = false;
} else if (Config.ParanoidTSO) {
AtomicTSOEmulationEnabled = true;
VectorAtomicTSOEmulationEnabled = true;
MemcpyAtomicTSOEmulationEnabled = true;
} else {
// Atomic TSO emulation only enabled if the config option is enabled.
AtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled;
@@ -376,12 +358,9 @@ private:
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, void* Ptr);
IR::AOTIRCaptureCache IRCaptureCache;
fextl::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool StartPaused = false;
bool IsMemoryShared = false;
bool SupportsHardwareTSO = false;
bool AtomicTSOEmulationEnabled = true;
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "FEXCore/Core/X86Enums.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
@@ -94,6 +93,7 @@ namespace x64 {
ARMEmitter::Reg::r20,
ARMEmitter::Reg::r21,
ARMEmitter::Reg::r22,
// PF/AF must be last.
REG_PF,
REG_AF,
};
@@ -610,7 +610,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
#endif
if (SetPredRegs) {
if (SetPredRegs && (EmitterCTX->HostFeatures.SupportsSVE256 || EmitterCTX->HostFeatures.SupportsSVE128)) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
@@ -622,6 +622,9 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
if (EmitterCTX->HostFeatures.SupportsSVE128) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
// Fill in the predicate register for the x87 ldst SVE optimization.
ptrue(ARMEmitter::SubRegSize::i16Bit, PRED_X87_SVEOPT, ARMEmitter::PredicatePattern::SVE_VL5);
}
}
@@ -1046,7 +1049,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
}
// Fill the static registers.
FillStaticRegs(true, PreserveSRAMask, PreserveSRAFPRMask);
FillStaticRegs(FPRs, PreserveSRAMask, PreserveSRAFPRMask);
// Pop the vector registers.
PopVectorRegisters(CanUseSVE256, DynamicFPRs);
@@ -18,10 +18,8 @@
#include <CodeEmitter/Emitter.h>
#include <CodeEmitter/Registers.h>
#include <array>
#include <cstddef>
#include <cstdint>
#include <utility>
#include <span>
namespace FEXCore::Context {
@@ -48,6 +46,10 @@ constexpr auto REG_AF = ARMEmitter::Reg::r27;
// Vector temporaries
constexpr auto VTMP1 = ARMEmitter::VReg::v0;
constexpr auto VTMP2 = ARMEmitter::VReg::v1;
// Predicate register for X87 SVE Optimization
constexpr auto SVE_OPT_PRED = ARMEmitter::PReg::p2;
#else
constexpr auto TMP1 = ARMEmitter::XReg::x10;
constexpr auto TMP2 = ARMEmitter::XReg::x11;
@@ -67,6 +69,9 @@ constexpr auto VTMP2 = ARMEmitter::VReg::v17;
constexpr auto EC_CALL_CHECKER_PC_REG = ARMEmitter::XReg::x9;
constexpr auto EC_ENTRY_CPUAREA_REG = ARMEmitter::XReg::x17;
// Predicate register for X87 SVE Optimization
constexpr auto SVE_OPT_PRED = ARMEmitter::PReg::p2;
// These structures are not included in the standard Windows headers, define the offsets of members we care about for EC here.
constexpr size_t TEB_CPU_AREA_OFFSET = 0x1788;
constexpr size_t TEB_PEB_OFFSET = 0x60;
@@ -74,8 +79,16 @@ constexpr size_t PEB_EC_CODE_BITMAP_OFFSET = 0x368;
constexpr size_t CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET = 0x1;
constexpr size_t CPU_AREA_EMULATOR_STACK_BASE_OFFSET = 0x8;
constexpr size_t CPU_AREA_EMULATOR_DATA_OFFSET = 0x30;
constexpr uint64_t EC_CODE_BITMAP_MAX_ADDRESS = 1ULL << 47;
#endif
// Will force one single instruction block to be generated first if set when entering the JIT filling SRA.
constexpr auto ENTRY_FILL_SRA_SINGLE_INST_REG = TMP1;
// Predicate to use in the X87 SVE optimization
constexpr ARMEmitter::PRegister PRED_X87_SVEOPT = ARMEmitter::PReg::p2;
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
// PRED_TMP_32B indicates a predicate register that indicates the first 32 bytes set to 1.
+14 -1
View File
@@ -39,6 +39,19 @@ namespace CPU {
{0xC90F'DAA2'2168'C235ULL, 0x0000'0000'0000'4000ULL}, // NAMED_VECTOR_X87_PI
{0x9A20'9A84'FBCF'F799ULL, 0x0000'0000'0000'3FFDULL}, // NAMED_VECTOR_X87_LOG10_2
{0xB172'17F7'D1CF'79ACULL, 0x0000'0000'0000'3FFEULL}, // NAMED_VECTOR_X87_LOG_2
{0x4F00'0000'4F00'0000ULL, 0x4F00'0000'4F00'0000ULL}, // NAMED_VECTOR_CVTMAX_F32_I32
{0x4F00'0000'4F00'0000ULL, 0x4F00'0000'4F00'0000ULL}, // NAMED_VECTOR_CVTMAX_F32_I32_UPPER
{0x5F00'0000'5F00'0000ULL, 0x5F00'0000'5F00'0000ULL}, // NAMED_VECTOR_CVTMAX_F32_I64
{0x41E0'0000'0000'0000ULL, 0x41E0'0000'0000'0000ULL}, // NAMED_VECTOR_CVTMAX_F64_I32
{0x41E0'0000'0000'0000ULL, 0x41E0'0000'0000'0000ULL}, // NAMED_VECTOR_CVTMAX_F64_I32_UPPER
{0x43E0'0000'0000'0000ULL, 0x43E0'0000'0000'0000ULL}, // NAMED_VECTOR_CVTMAX_F64_I64
{0x8000'0000'8000'0000ULL, 0x8000'0000'8000'0000ULL}, // NAMED_VECTOR_CVTMAX_I32
{0x8000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_CVTMAX_I64
{0x0000'0000'0000'0000ULL, 0x0000'0000'0000'8000ULL}, // NAMED_VECTOR_F80_SIGN_MASK
{0x5A82'7999'5A82'7999ULL, 0x5A82'7999'5A82'7999ULL}, // NAMED_VECTOR_SHA1RNDS_K0
{0x6ED9'EBA1'6ED9'EBA1ULL, 0x6ED9'EBA1'6ED9'EBA1ULL}, // NAMED_VECTOR_SHA1RNDS_K1
{0x8F1B'BCDC'8F1B'BCDCULL, 0x8F1B'BCDC'8F1B'BCDCULL}, // NAMED_VECTOR_SHA1RNDS_K2
{0xCA62'C1D6'CA62'C1D6ULL, 0xCA62'C1D6'CA62'C1D6ULL}, // NAMED_VECTOR_SHA1RNDS_K3
};
constexpr static auto PSHUFLW_LUT {[]() consteval {
@@ -364,7 +377,7 @@ namespace CPU {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Buffer.Size, true));
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
+12 -18
View File
@@ -80,9 +80,13 @@ namespace CPU {
struct JITCodeTail {
// The total size of the codeblock from [BlockBegin, BlockBegin+Size).
size_t Size;
// RIP that the block's entry comes from.
uint64_t RIP;
// The length of the guest code for this block.
size_t GuestSize;
// Number of RIP entries for this JIT Code section.
uint32_t NumberOfRIPEntries;
@@ -92,23 +96,10 @@ namespace CPU {
// Shared-code modification spin-loop futex.
uint32_t SpinLockFutex;
uint32_t _Pad;
};
// If this block represents a single guest instruction.
bool SingleInst;
// Entries that live after the JITCodeTail.
// These entries correlate JIT code regions with guest RIP regions.
// Using these entries FEX is able to reconstruct the guest RIP accurately when an instruction cause a signal fault.
// Packed using 16-bit entries to ensure the size isn't too large.
// These smaller sizes means that each entry is relative to each other instead of absolute offset from the start of the JIT block.
// When reconstructing the RIP, each entry must be walked linearly and accumulated with the previous entries.
// This is a trade-off between compression inside the JIT code space and execution time when reconstruction the RIP.
// RIP reconstruction when faulting is less likely so we are requiring the accumulation.
struct JITRIPReconstructEntries {
// The Host PC offset from the previous entry.
uint16_t HostPCOffset;
// How much to offset the RIP from the previous entry.
uint16_t GuestRIPOffset;
uint8_t _Pad[3];
};
/**
@@ -119,14 +110,17 @@ namespace CPU {
*
* This is a thread specific compilation unit since there is one CPUBackend per guest thread
*
* @param Size - The byte size of the guest code for this block
* @param SingleInst - If this block represents a single guest instruction
* @param IR - IR that maps to the IR for this RIP
* @param DebugData - Debug data that is available for this IR indirectly
* @param CheckTF - If EFLAGS.TF checks should be emitted at the start of the block
*
* @return Information about the compiled code block.
*/
[[nodiscard]]
virtual CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData) = 0;
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, const FEXCore::IR::RegisterAllocationData* RAData, bool CheckTF) = 0;
/**
* @brief Relocates a block of code from the JIT code object cache
+18 -5
View File
@@ -90,7 +90,7 @@ namespace ProductNames {
#endif
} // namespace ProductNames
static uint32_t GetCPUID() {
uint32_t GetCPUID_Syscall() {
uint32_t CPU {};
FHU::Syscalls::getcpu(&CPU, nullptr);
return CPU;
@@ -138,6 +138,12 @@ uint32_t GetCycleCounterFrequency() {
return Result;
}
uint32_t GetCPUID_TPIDRRO() {
uint64_t Result {};
__asm("mrs %[Res], TPIDRRO_EL0" : [Res] "=r"(Result));
return Result;
}
void CPUIDEmu::SetupHostHybridFlag() {
PerCPUData.resize(Cores);
@@ -186,7 +192,6 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x41, 0xd4e, 1, ProductNames::ARM_X3}, // X3
{0x41, 0xd4d, 1, ProductNames::ARM_A715}, // A715
{0x41, 0xd4f, 1, ProductNames::ARM_V2}, // V2
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd4b, 1, ProductNames::ARM_A78C}, // A78C
{0x41, 0xd4a, 1, ProductNames::ARM_E1}, // E1
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
@@ -895,11 +900,11 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) con
// Extended processor and feature bits
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) const {
// RDTSCP is disabled on WIN32/Wine because there is no sane way to query processor ID.
#ifndef _WIN32
constexpr uint32_t SUPPORTS_RDTSCP = 1;
#else
constexpr uint32_t SUPPORTS_RDTSCP = 0;
// RDTSCP under WIN32 is only supported if CPUIndex is available in TPIDRRO.
const uint32_t SUPPORTS_RDTSCP = SupportsCPUIndexInTPIDRRO;
#endif
FEXCore::CPUID::FunctionResults Res {};
@@ -1213,12 +1218,20 @@ FEXCore::CPUID::XCRResults CPUIDEmu::XCRFunction_0h() const {
}
CPUIDEmu::CPUIDEmu(const FEXCore::Context::ContextImpl* ctx)
: CTX {ctx} {
: CTX {ctx}
, SupportsCPUIndexInTPIDRRO {CTX->HostFeatures.SupportsCPUIndexInTPIDRRO}
, GetCPUID {GetCPUID_Syscall} {
Cores = CTX->HostFeatures.CPUMIDRs.size();
// Setup some state tracking
SetupHostHybridFlag();
SetupFeatures();
#ifdef _M_ARM_64
if (SupportsCPUIndexInTPIDRRO) {
GetCPUID = GetCPUID_TPIDRRO;
}
#endif
}
} // namespace FEXCore
+4
View File
@@ -115,6 +115,7 @@ public:
private:
const FEXCore::Context::ContextImpl* CTX;
bool SupportsCPUIndexInTPIDRRO {};
bool Hybrid {};
uint32_t Cores {};
FEX_CONFIG_OPT(HideHypervisorBit, HIDEHYPERVISORBIT);
@@ -510,5 +511,8 @@ private:
// 0x8000'001F: AMD Secure Encryption
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
}};
using GetCPUIDPtr = uint32_t (*)();
GetCPUIDPtr GetCPUID;
};
} // namespace FEXCore
+127 -192
View File
@@ -9,14 +9,14 @@ $end_info$
*/
#include <cstdint>
#include "Interface/Core/ArchHelpers//Arm64Emitter.h"
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/JIT/JITCore.h"
#include "Interface/Core/JIT/JITClass.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/IR/IR.h"
@@ -28,6 +28,7 @@ $end_info$
#include "Utils/Allocator.h"
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/SpinWaitLock.h"
#include "Utils/variable_length_integer.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -112,31 +113,59 @@ ContextImpl::~ContextImpl() {
}
}
uint64_t ContextImpl::RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) {
const auto Frame = Thread->CurrentFrame;
struct GetFrameBlockInfoResult {
const CPU::CPUBackend::JITCodeHeader* InlineHeader;
const CPU::CPUBackend::JITCodeTail* InlineTail;
};
static GetFrameBlockInfoResult GetFrameBlockInfo(FEXCore::Core::CpuStateFrame* Frame) {
const uint64_t BlockBegin = Frame->State.InlineJITBlockHeader;
auto InlineHeader = reinterpret_cast<const CPU::CPUBackend::JITCodeHeader*>(BlockBegin);
if (InlineHeader) {
auto InlineTail = reinterpret_cast<const CPU::CPUBackend::JITCodeTail*>(Frame->State.InlineJITBlockHeader + InlineHeader->OffsetToBlockTail);
auto RIPEntries = reinterpret_cast<const CPU::CPUBackend::JITRIPReconstructEntries*>(
Frame->State.InlineJITBlockHeader + InlineHeader->OffsetToBlockTail + InlineTail->OffsetToRIPEntries);
return {InlineHeader, InlineTail};
}
return {InlineHeader, nullptr};
}
bool ContextImpl::IsAddressInCurrentBlock(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, uint64_t Size) {
auto [_, InlineTail] = GetFrameBlockInfo(Thread->CurrentFrame);
return InlineTail && (Address + Size > InlineTail->RIP && Address < InlineTail->RIP + InlineTail->GuestSize);
}
bool ContextImpl::IsCurrentBlockSingleInst(FEXCore::Core::InternalThreadState* Thread) {
auto [_, InlineTail] = GetFrameBlockInfo(Thread->CurrentFrame);
return InlineTail && InlineTail->SingleInst;
}
uint64_t ContextImpl::RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) {
const auto Frame = Thread->CurrentFrame;
const uint64_t BlockBegin = Frame->State.InlineJITBlockHeader;
auto [InlineHeader, InlineTail] = GetFrameBlockInfo(Thread->CurrentFrame);
if (InlineHeader) {
// Check if the host PC is currently within a code block.
// If it is then RIP can be reconstructed from the beginning of the code block.
// This is currently as close as FEX can get RIP reconstructions.
if (HostPC >= reinterpret_cast<uint64_t>(BlockBegin) && HostPC < reinterpret_cast<uint64_t>(BlockBegin + InlineTail->Size)) {
auto RIPEntry =
reinterpret_cast<const uint8_t*>(Frame->State.InlineJITBlockHeader + InlineHeader->OffsetToBlockTail + InlineTail->OffsetToRIPEntries);
// Reconstruct RIP from JIT entries for this block.
uint64_t StartingHostPC = BlockBegin;
uint64_t StartingGuestRIP = InlineTail->RIP;
for (uint32_t i = 0; i < InlineTail->NumberOfRIPEntries; ++i) {
const auto& RIPEntry = RIPEntries[i];
if (HostPC >= (StartingHostPC + RIPEntry.HostPCOffset)) {
auto HostPCOffset = FEXCore::Utils::vl64::Decode(RIPEntry);
RIPEntry += HostPCOffset.Size;
auto GuestRIPOffset = FEXCore::Utils::vl64::Decode(RIPEntry);
RIPEntry += GuestRIPOffset.Size;
if (HostPC >= (StartingHostPC + HostPCOffset.Integer)) {
// We are beyond this entry, keep going forward.
StartingHostPC += RIPEntry.HostPCOffset;
StartingGuestRIP += RIPEntry.GuestRIPOffset;
StartingHostPC += HostPCOffset.Integer;
StartingGuestRIP += GuestRIPOffset.Integer;
} else {
// Passed where the Host PC is at. Break now.
break;
@@ -150,7 +179,8 @@ uint64_t ContextImpl::RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* T
return Frame->State.rip;
}
uint32_t ContextImpl::ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, uint64_t* HostGPRs, uint64_t PSTATE) {
uint32_t ContextImpl::ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, const uint64_t* HostGPRs,
uint64_t PSTATE) {
const auto Frame = Thread->CurrentFrame;
uint32_t EFLAGS {};
@@ -160,6 +190,7 @@ uint32_t ContextImpl::ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadSt
case X86State::RFLAG_CF_RAW_LOC:
case X86State::RFLAG_PF_RAW_LOC:
case X86State::RFLAG_AF_RAW_LOC:
case X86State::RFLAG_TF_RAW_LOC:
case X86State::RFLAG_ZF_RAW_LOC:
case X86State::RFLAG_SF_RAW_LOC:
case X86State::RFLAG_OF_RAW_LOC:
@@ -212,6 +243,9 @@ uint32_t ContextImpl::ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadSt
uint32_t AF = ((Frame->State.af_raw ^ PFByte) & (1 << 4)) ? 1 : 0;
EFLAGS |= AF << X86State::RFLAG_AF_RAW_LOC;
uint8_t TFByte = Frame->State.flags[X86State::RFLAG_TF_RAW_LOC];
EFLAGS |= (TFByte & 1) << X86State::RFLAG_TF_RAW_LOC;
// DF is pretransformed, undo the transform from 1/-1 back to 0/1
uint8_t DFByte = Frame->State.flags[X86State::RFLAG_DF_RAW_LOC];
if (DFByte & 0x80) {
@@ -354,8 +388,6 @@ bool ContextImpl::InitCore() {
if (Config.GdbServer) {
// If gdbserver is enabled then this needs to be enabled.
Config.NeedsPendingInterruptFaultCheck = true;
// FEX needs to start paused when gdb is enabled.
StartPaused = true;
}
return true;
@@ -365,29 +397,17 @@ void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState* Thread, uin
static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->ExecuteJITCallback(Thread->CurrentFrame, RIP);
}
FEXCore::Context::ExitReason ContextImpl::RunUntilExit(FEXCore::Core::InternalThreadState* Thread) {
ExecutionThread(Thread);
CoreShuttingDown.store(true);
if (CustomExitHandler) {
CustomExitHandler(Thread, FEXCore::Context::ExitReason::EXIT_SHUTDOWN);
return Thread->ExitReason;
}
return FEXCore::Context::ExitReason::EXIT_SHUTDOWN;
}
void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
}
if (CodeObjectCacheService) {
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
void ContextImpl::InitializeThreadTLSData(FEXCore::Core::InternalThreadState* Thread) {
// Let's do some initial bookkeeping here
#ifndef _WIN32
Alloc::OSAllocator::RegisterTLSData(Thread);
#endif
// If it is the parent thread that died then just leave
FEX_TODO("This doesn't make sense when the parent thread doesn't outlive its children");
}
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
@@ -402,8 +422,6 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Dispatcher->InitThreadPointers(Thread);
Thread->CTX = this;
Thread->PassManager->AddDefaultPasses(this);
Thread->PassManager->AddDefaultValidationPasses();
@@ -418,7 +436,9 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
FEXCore::Core::InternalThreadState*
ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXCore::Core::CPUState* NewThreadState, uint64_t ParentTID) {
FEXCore::Core::InternalThreadState* Thread = new FEXCore::Core::InternalThreadState {};
FEXCore::Core::InternalThreadState* Thread = new FEXCore::Core::InternalThreadState {
.CTX = this,
};
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = StackPointer;
Thread->CurrentFrame->State.rip = InitialRIP;
@@ -443,13 +463,7 @@ ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXC
return Thread;
}
void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState* Thread, bool NeedsTLSUninstall) {
if (NeedsTLSUninstall) {
#ifndef _WIN32
Alloc::OSAllocator::UninstallTLSData(Thread);
#endif
}
void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState* Thread) {
FEXCore::Allocator::VirtualProtect(&Thread->InterruptFaultPage, sizeof(Thread->InterruptFaultPage),
Allocator::ProtectOptions::Read | Allocator::ProtectOptions::Write);
delete Thread;
@@ -459,6 +473,7 @@ void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState* Thread, bool
void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread, bool Child) {
Allocator::UnlockAfterFork(LiveThread, Child);
Profiler::PostForkAction(Child);
if (Child) {
CodeInvalidationMutex.StealAndDropActiveLocks();
if (Config.StrictInProcessSplitLocks) {
@@ -482,14 +497,10 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
}
#endif
void ContextImpl::AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, void* Ptr) {
Thread->LookupCache->AddBlockMapping(Address, Ptr);
}
void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread) {
FEXCORE_PROFILE_INSTANT("ClearCodeCache");
{
if (CodeObjectCacheService) {
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
@@ -508,40 +519,6 @@ static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter*
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
};
// IRStorageBase with fully owned memory
struct IRListCopy : public IR::IRStorageBase {
std::span<std::byte> IRData;
std::span<std::byte> ListData;
// TODO: Consider defaulting to empty RAData instead?
IR::RegisterAllocationData::UniquePtr RADataInternal;
IRListCopy(const IR::IRListView& view, IR::RegisterAllocationData::UniquePtr RAData)
: RADataInternal(std::move(RAData)) {
std::byte* Storage = reinterpret_cast<std::byte*>(FEXCore::Allocator::malloc(view.GetDataSize() + view.GetListSize()));
IRData = {Storage, Storage + view.GetDataSize()};
ListData = {Storage + view.GetDataSize(), Storage + view.GetDataSize() + view.GetListSize()};
memcpy(IRData.data(), (char*)view.GetData(), IRData.size());
memcpy(ListData.data(), (char*)view.GetListData(), ListData.size());
}
IRListCopy(const IRListCopy& other) = delete;
IRListCopy(IRListCopy&& other) = delete;
~IRListCopy() {
FEXCore::Allocator::free(IRData.data());
}
const IR::RegisterAllocationData* RAData() override {
return RADataInternal.get();
}
IR::IRListView GetIRView() override {
return IR::IRListView {IRData.data(), ListData.data(), IRData.size(), ListData.size()};
}
};
ContextImpl::GenerateIRResult
ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
@@ -570,6 +547,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
GuestCode = reinterpret_cast<const uint8_t*>(GuestRIP);
bool HadDispatchError {false};
bool HadInvalidInst {false};
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP, MaxInst,
[Thread](uint64_t BlockEntry, uint64_t Start, uint64_t Length) {
@@ -583,7 +561,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount);
const uint8_t GPRSize = GetGPRSize();
const auto GPRSize = GetGPROpSize();
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
const FEXCore::Frontend::Decoder::DecodedBlocks& Block = CodeBlocks->at(j);
@@ -599,7 +577,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (InstsInBlock == 0) {
// Special case for an empty instruction block.
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, Block.Entry - GuestRIP));
}
for (size_t i = 0; i < InstsInBlock; ++i) {
@@ -642,8 +620,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->ExitFunction(
Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry + BlockInstructionsLength - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, Block.Entry + BlockInstructionsLength - GuestRIP));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
@@ -668,30 +645,34 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
++TotalInstructions;
}
} else {
if (TableInfo) {
LogMan::Msg::EFmt("Invalid or Unknown instruction: {} 0x{:x}", TableInfo->Name ?: "UND", Block.Entry - GuestRIP);
}
// Invalid instruction
Thread->OpDispatcher->InvalidOp(DecodedInfo);
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry - GuestRIP));
if (!BlockInstructionsLength) {
// SMC can modify block contents and patch invalid instructions to valid ones inline.
// End blocks upon encountering them and only emit an invalid opcode exception if there are no prior instructions in the block (that could have modified it to be valid).
if (TableInfo) {
LogMan::Msg::EFmt("Invalid or Unknown instruction: {} 0x{:x}", TableInfo->Name ?: "UND", Block.Entry - GuestRIP);
}
Thread->OpDispatcher->InvalidOp(DecodedInfo);
}
HadInvalidInst = true;
}
const bool NeedsBlockEnd =
(HadDispatchError && TotalInstructions > 0) || (Thread->OpDispatcher->NeedsBlockEnder() && i + 1 == InstsInBlock);
const bool NeedsBlockEnd = (HadDispatchError && TotalInstructions > 0) ||
(Thread->OpDispatcher->NeedsBlockEnder() && i + 1 == InstsInBlock) || HadInvalidInst;
// If we had a dispatch error then leave early
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return {nullptr, 0, 0, 0, 0};
return {{}, nullptr, 0, 0, 0, 0};
}
if (NeedsBlockEnd) {
const uint8_t GPRSize = GetGPRSize();
// We had some instructions. Early exit
Thread->OpDispatcher->ExitFunction(
Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry + BlockInstructionsLength - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, Block.Entry + BlockInstructionsLength - GuestRIP));
break;
}
@@ -718,19 +699,16 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(IREmitter);
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr;
// Debug
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP,
Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
IRDumper(Thread, IREmitter, GuestRIP, RAData);
}
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
auto IRList = fextl::make_unique<IRListCopy>(IREmitter->ViewIR(), std::move(RAData));
IREmitter->DelayedDisownBuffer();
return {
.IR = std::move(IRList),
.IRView = IREmitter->ViewIR(),
.RAData = RAData,
.TotalInstructions = TotalInstructions,
.TotalInstructionsLength = TotalInstructionsLength,
.StartAddr = Thread->FrontendDecoder->DecodedMinAddress,
@@ -747,9 +725,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (CompiledCode) {
return {
.CompiledCode = CompiledCode,
.IR = nullptr, // No IR/RA data generated
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.GeneratedIR = false, // nullptr here ensures IR cache mechanisms won't run
.StartAddr = 0, // Unused
.Length = 0, // Unused
};
@@ -764,55 +740,39 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
}
}
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
FEXCore::Core::DebugData* DebugData {};
uint64_t StartAddr {};
uint64_t Length {};
// AOT IR bookkeeping and cache
{
auto IRFromAOT = IRCaptureCache.PreGenerateIRFetch(Thread, GuestRIP);
if (IRFromAOT) {
// Setup pointers to internal structures
IR = std::move(IRFromAOT->IR);
DebugData = IRFromAOT->DebugData;
StartAddr = IRFromAOT->StartAddr;
Length = IRFromAOT->Length;
}
// Generate IR + Meta Info
auto [IRView, RAData, TotalInstructions, TotalInstructionsLength, StartAddr, Length] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
return {nullptr, nullptr, 0, 0};
}
auto DebugData = fextl::make_unique<FEXCore::Core::DebugData>();
if (!IR) {
// Generate IR + Meta Info
auto [IRCopy, TotalInstructions, TotalInstructionsLength, _StartAddr, _Length] = GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
// If the trap flag is set we generate single instruction blocks that each check to generate a single step exception.
bool TFSet = Thread->CurrentFrame->State.flags[X86State::RFLAG_TF_RAW_LOC];
// Setup pointers to internal structures
IR = std::move(IRCopy);
DebugData = new FEXCore::Core::DebugData();
StartAddr = _StartAddr;
Length = _Length;
}
if (!IR) {
return {};
}
// Attempt to get the CPU backend to compile this code
auto IRView = IR->GetIRView();
auto CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, Length, TotalInstructions == 1, &*IRView, DebugData.get(), RAData, TFSet);
// Release the IR
Thread->OpDispatcher->DelayedDisownBuffer();
return {
// FEX currently throws away the CPUBackend::CompiledCode object other than the entrypoint
// In the future with code caching getting wired up, we will pass the rest of the data forward.
// TODO: Pass the data forward when code caching is wired up to this.
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, &IRView, DebugData, IR->RAData()).BlockEntry,
.IR = std::move(IR),
.DebugData = DebugData,
.GeneratedIR = true,
.CompiledCode = CompiledCode.BlockEntry,
.DebugData = std::move(DebugData),
.StartAddr = StartAddr,
.Length = Length,
};
}
uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("CompileBlock");
auto Thread = Frame->Thread;
FEXCORE_PROFILE_SCOPED("CompileBlock");
FEXCORE_PROFILE_ACCUMULATION(Thread, AccumulatedJITTime);
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
auto lk = GuardSignalDeferringSection<std::shared_lock>(CodeInvalidationMutex, Thread);
@@ -823,7 +783,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
return HostCode;
}
auto [CodePtr, IR, DebugData, GeneratedIR, StartAddr, Length] = CompileCode(Thread, GuestRIP, MaxInst);
auto [CodePtr, DebugData, StartAddr, Length] = CompileCode(Thread, GuestRIP, MaxInst);
if (CodePtr == nullptr) {
return 0;
}
@@ -874,55 +834,34 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
// Clear any relocations that might have been generated
Thread->CPUBackend->ClearRelocations();
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, std::move(IR), DebugData, GeneratedIR)) {
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, {}, DebugData.get(), false)) {
// Early exit
return (uintptr_t)CodePtr;
}
// Insert to lookup cache
// Pages containing this block are added via AddBlockExecutableRange before each page gets accessed in the frontend
AddBlockMapping(Thread, GuestRIP, CodePtr);
Thread->LookupCache->AddBlockMapping(GuestRIP, CodePtr);
return (uintptr_t)CodePtr;
}
void ContextImpl::ExecutionThread(FEXCore::Core::InternalThreadState* Thread) {
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_WAITING;
uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
FEXCORE_PROFILE_SCOPED("CompileSingleStep");
auto Thread = Frame->Thread;
InitializeThreadTLSData(Thread);
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
auto lk = GuardSignalDeferringSection<std::shared_lock>(CodeInvalidationMutex, Thread);
// Now notify the thread that we are initialized
Thread->ThreadWaiting.NotifyAll();
if (StartPaused || Thread->StartPaused) {
// Parent thread doesn't need to wait to run
Thread->StartRunning.Wait();
auto [CodePtr, DebugData, StartAddr, Length] = CompileCode(Thread, GuestRIP, 1);
if (CodePtr == nullptr) {
return 0;
}
if (!Thread->RunningEvents.EarlyExit.load()) {
Thread->RunningEvents.WaitingToStart = false;
// Clear any relocations that might have been generated
Thread->CPUBackend->ClearRelocations();
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_NONE;
Thread->RunningEvents.Running = true;
static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = false;
}
{
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
// If it is the parent thread that died then just leave
FEX_TODO("This doesn't make sense when the parent thread doesn't outlive its children");
#ifndef _WIN32
Alloc::OSAllocator::UninstallTLSData(Thread);
#endif
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
@@ -950,6 +889,10 @@ void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* T
}
void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) {
if (!Thread) {
return;
}
if (!IsMemoryShared) {
IsMemoryShared = true;
UpdateAtomicTSOEmulationConfig();
@@ -962,19 +905,10 @@ void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) {
}
}
void ContextImpl::ThreadAddBlockLink(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestDestination,
FEXCore::Context::ExitFunctionLinkData* HostLink, const FEXCore::Context::BlockDelinkerFunc& delinker) {
auto lk = GuardSignalDeferringSection<std::shared_lock>(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
Thread->LookupCache->AddBlockLink(GuestDestination, HostLink, delinker);
}
void ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
}
@@ -996,11 +930,11 @@ ContextImpl::AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandl
}
void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t GuestThunkEntrypoint) {
LOGMAN_THROW_AA_FMT(Entrypoint, "Tried to link null pointer address to guest function");
LOGMAN_THROW_AA_FMT(GuestThunkEntrypoint, "Tried to link address to null pointer guest function");
LOGMAN_THROW_A_FMT(Entrypoint, "Tried to link null pointer address to guest function");
LOGMAN_THROW_A_FMT(GuestThunkEntrypoint, "Tried to link address to null pointer guest function");
if (!Config.Is64BitMode) {
LOGMAN_THROW_AA_FMT((Entrypoint >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_AA_FMT((GuestThunkEntrypoint >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_A_FMT((Entrypoint >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_A_FMT((GuestThunkEntrypoint >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
}
LogMan::Msg::DFmt("Thunks: Adding guest trampoline from address {:#x} to guest function {:#x}", Entrypoint, GuestThunkEntrypoint);
@@ -1013,12 +947,13 @@ void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t Gu
IRHeader.first->Blocks = emit->WrapNode(Block);
emit->SetCurrentCodeBlock(Block);
const uint8_t GPRSize = GetGPRSize();
const auto GPRSize = GetGPROpSize();
if (GPRSize == 8) {
if (GPRSize == IR::OpSize::i64Bit) {
emit->_StoreRegister(emit->_Constant(Entrypoint), X86State::REG_R11, IR::GPRClass, GPRSize);
} else {
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(8, 8, emit->_Constant(Entrypoint)), offsetof(Core::CPUState, mm[0][0]));
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->_Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
}
emit->_ExitFunction(emit->_Constant(GuestThunkEntrypoint));
},
@@ -46,6 +46,8 @@ Dispatcher::~Dispatcher() {
}
void Dispatcher::EmitDispatcher() {
// Don't modify TMP3 since it contains our RIP once the block doesn't exist
auto RipReg = TMP3;
#ifdef VIXL_DISASSEMBLER
const auto DisasmBegin = GetCursorAddress<const vixl::aarch64::Instruction*>();
#endif
@@ -61,8 +63,9 @@ void Dispatcher::EmitDispatcher() {
// }
ARMEmitter::ForwardLabel l_CTX;
ARMEmitter::SingleUseForwardLabel l_Sleep;
ARMEmitter::SingleUseForwardLabel l_CompileBlock;
ARMEmitter::ForwardLabel l_Sleep;
ARMEmitter::ForwardLabel l_CompileBlock;
ARMEmitter::ForwardLabel l_CompileSingleStep;
// Push all the register we need to save
PushCalleeSavedRegisters();
@@ -81,6 +84,7 @@ void Dispatcher::EmitDispatcher() {
FillStaticRegs();
ARMEmitter::BiDirectionalLabel LoopTop {};
ARMEmitter::ForwardLabel CompileSingleStep;
#ifdef _M_ARM_64EC
b(&LoopTop);
@@ -89,6 +93,10 @@ void Dispatcher::EmitDispatcher() {
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
FillStaticRegs();
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
// Force a single instruction block if ENTRY_FILL_SRA_SINGLE_INST_REG is nonzero entering the JIT, used for inline SMC handling.
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
// Enter JIT
b(&LoopTop);
@@ -116,10 +124,11 @@ void Dispatcher::EmitDispatcher() {
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
// Load in our RIP
// Don't modify TMP3 since it contains our RIP once the block doesn't exist
auto RipReg = TMP3;
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbnz(ARMEmitter::Size::i32Bit, TMP1, &CompileSingleStep);
// L1 Cache
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
@@ -204,37 +213,21 @@ void Dispatcher::EmitDispatcher() {
ret();
}
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
SpillStaticRegs(TMP1);
// Clobbers TMP1/2
auto EmitSignalGuardedRegion = [&](auto Body) {
#ifndef _WIN32
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
str(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#endif
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 1);
strb(TMP1.W(), TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
mov(ARMEmitter::XReg::x0, STATE);
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillStaticRegs();
Body();
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
@@ -250,17 +243,38 @@ void Dispatcher::EmitDispatcher() {
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
#endif
};
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
mov(ARMEmitter::XReg::x0, STATE);
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillStaticRegs();
});
br(TMP1);
}
// Need to create the block
{
Bind(&NoBlock);
#ifdef _M_ARM_64EC
// Clobbers TMP1/2
auto EmitECExitCheck = [&]() {
// Check the EC code bitmap incase we need to exit the JIT to call into native code.
ARMEmitter::SingleUseForwardLabel l_NotECCode;
ARMEmitter::ForwardLabel l_NotECCode;
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
@@ -277,56 +291,83 @@ void Dispatcher::EmitDispatcher() {
br(TMP2);
Bind(&l_NotECCode);
};
#endif
SpillStaticRegs(TMP1);
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x2, RipReg);
}
#ifndef _WIN32
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#endif
// Need to create the block
{
Bind(&NoBlock);
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
EmitECExitCheck();
#endif
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
// x2 contains guest RIP
mov(ARMEmitter::XReg::x3, 0);
ldr(ARMEmitter::XReg::x4, &l_CompileBlock);
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4); // { CTX, Frame, RIP, MaxInst }
}
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x2, RipReg);
}
FillStaticRegs();
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
// x2 contains guest RIP
mov(ARMEmitter::XReg::x3, 0);
ldr(ARMEmitter::XReg::x4, &l_CompileBlock);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4); // { CTX, Frame, RIP, MaxInst }
}
// Result is now in x0
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillStaticRegs();
});
// Jump to the compiled block
br(TMP1);
}
{
Bind(&CompileSingleStep);
#ifdef _M_ARM_64EC
ldr(TMP1, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP1, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
EmitECExitCheck();
#endif
#ifndef _WIN32
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
// Trigger segfault if any deferred signals are pending
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
#endif
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x2, RipReg);
}
b(&LoopTop);
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
// x2 contains guest RIP
ldr(ARMEmitter::XReg::x4, &l_CompileSingleStep);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4); // { CTX, Frame, RIP }
}
// Result is now in x0
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillStaticRegs();
});
// Jump to the compiled block
br(TMP1);
}
{
@@ -505,8 +546,11 @@ void Dispatcher::EmitDispatcher() {
Bind(&l_Sleep);
dc64(reinterpret_cast<uint64_t>(SleepThread));
Bind(&l_CompileBlock);
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::Context::ContextImpl::CompileBlock);
dc64(PMF.GetConvertedPointer());
FEXCore::Utils::MemberFunctionToPointerCast PMFCompileBlock(&FEXCore::Context::ContextImpl::CompileBlock);
dc64(PMFCompileBlock.GetConvertedPointer());
Bind(&l_CompileSingleStep);
FEXCore::Utils::MemberFunctionToPointerCast PMFCompileSingleStep(&FEXCore::Context::ContextImpl::CompileSingleStep);
dc64(PMFCompileSingleStep.GetConvertedPointer());
Start = reinterpret_cast<uint64_t>(DispatchPtr);
End = GetCursorAddress<uint64_t>();
+170 -88
View File
@@ -75,7 +75,7 @@ Decoder::~Decoder() {
uint8_t Decoder::ReadByte() {
uint8_t Byte = InstStream[InstructionSize];
LOGMAN_THROW_AA_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
Instruction[InstructionSize] = Byte;
InstructionSize++;
return Byte;
@@ -87,7 +87,7 @@ uint8_t Decoder::PeekByte(uint8_t Offset) const {
}
uint64_t Decoder::ReadData(uint8_t Size) {
LOGMAN_THROW_AA_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
LOGMAN_THROW_A_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
std::memcpy(&Res, &InstStream[InstructionSize], Size);
@@ -220,7 +220,8 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
// The invalid encoding types are described at Table 1-12. "promoted nsigned is always non-zero"
{
// If we have a VSIB byte (as opposed to SIB), then the index register is a vector.
const bool IsIndexVector = (DecodeInst->TableInfo->Flags & InstFlags::FLAGS_VEX_VSIB) != 0;
// DecodeInst->TableInfo may be null in the case of 3DNow! ModRM decoding.
const bool IsIndexVector = DecodeInst->TableInfo && (DecodeInst->TableInfo->Flags & InstFlags::FLAGS_VEX_VSIB) != 0;
uint8_t InvalidSIBIndex = 0b100; ///< SIB Index where there is no register encoding.
if (IsIndexVector) {
DecodeInst->Flags |= X86Tables::DecodeFlags::FLAG_VSIB_BYTE;
@@ -234,7 +235,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
Operand->Data.SIB.Base = MapModRMToReg(BaseREX, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
}
LOGMAN_THROW_AA_FMT(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
LOGMAN_THROW_A_FMT(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
if (Displacement) {
uint64_t Literal = ReadData(Displacement);
@@ -281,10 +282,10 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
return false;
}
LOGMAN_THROW_AA_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P), "Group Ops "
"should have "
"been decoded "
"before this!");
LOGMAN_THROW_A_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P), "Group Ops "
"should have "
"been decoded "
"before this!");
uint8_t DestSize {};
const bool HasWideningDisplacement =
@@ -403,7 +404,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_RAX) ? FEXCore::X86State::REG_RAX : FEXCore::X86State::REG_RDX;
CurrentDest = &DecodeInst->Src[0];
} else if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_REX_IN_BYTE)) {
LOGMAN_THROW_AA_FMT(!HasMODRM, "This instruction shouldn't have ModRM!");
LOGMAN_THROW_A_FMT(!HasMODRM, "This instruction shouldn't have ModRM!");
// If the REX is in the byte that means the lower nibble of the OP contains the destination GPR
// This also means that the destination is always a GPR on these ones
@@ -521,7 +522,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
}
if (Bytes != 0) {
LOGMAN_THROW_AA_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
LOGMAN_THROW_A_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
DecodeInst->Src[CurrentSrc].Data.Literal.Size = Bytes;
@@ -544,8 +545,8 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
LOGMAN_THROW_AA_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining", DecodeInst->PC,
DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining", DecodeInst->PC,
DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
}
@@ -562,7 +563,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
return false;
}
LOGMAN_THROW_AA_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
LOGMAN_THROW_A_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
// A normal instruction is the most likely.
if (Info->Type == FEXCore::X86Tables::TYPE_INST) [[likely]] {
@@ -606,13 +607,13 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
uint16_t LocalOp = OPD(Info->Type, PrefixType, ModRM.reg);
FEXCore::X86Tables::X86InstInfo* LocalInfo = &SecondInstGroupOps[LocalOp];
#undef OPD
if (LocalInfo->Type == FEXCore::X86Tables::TYPE_SECOND_GROUP_MODRM) {
if (LocalInfo->Type == FEXCore::X86Tables::TYPE_SECOND_GROUP_MODRM && ModRM.mod == 0b11) {
// Everything in this group is privileged instructions aside from XGETBV
constexpr std::array<uint8_t, 8> RegToField = {
255, 0, 1, 2, 255, 255, 255, 3,
};
uint8_t Field = RegToField[ModRM.reg];
LOGMAN_THROW_AA_FMT(Field != 255, "Invalid field selected!");
LOGMAN_THROW_A_FMT(Field != 255, "Invalid field selected!");
LocalOp = (Field << 3) | ModRM.rm;
return NormalOp(&SecondModRMTableOps[LocalOp], LocalOp);
@@ -926,8 +927,9 @@ void Decoder::BranchTargetInMultiblockRange() {
// If the RIP setting is conditional AND within our symbol range then it can be considered for multiblock
uint64_t TargetRIP = 0;
const uint8_t GPRSize = CTX->GetGPRSize();
const auto GPRSize = CTX->GetGPROpSize();
bool Conditional = true;
const auto InstEnd = DecodeInst->PC + DecodeInst->InstSize;
switch (DecodeInst->OP) {
case 0x70 ... 0x7F: // Conditional JUMP
@@ -936,17 +938,17 @@ void Decoder::BranchTargetInMultiblockRange() {
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
// Target offset is PC + InstSize + Literal
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
TargetRIP = InstEnd + DecodeInst->Src[0].Literal();
break;
}
case 0xE9:
case 0xEB: // Both are unconditional JMP instructions
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
TargetRIP = InstEnd + DecodeInst->Src[0].Literal();
Conditional = false;
break;
case 0xE8: // Call - Immediate target, We don't want to inline calls
if (ExternalBranches) {
ExternalBranches->insert(DecodeInst->PC + DecodeInst->InstSize);
ExternalBranches->insert(InstEnd);
}
[[fallthrough]];
case 0xC2: // RET imm
@@ -954,13 +956,15 @@ void Decoder::BranchTargetInMultiblockRange() {
default: return; break;
}
if (GPRSize == 4) {
if (GPRSize == IR::OpSize::i32Bit) {
// If we are running a 32bit guest then wrap around addresses that go above 32bit
TargetRIP &= 0xFFFFFFFFU;
}
// If the target RIP is x86 code within the symbol ranges then we are golden
bool ValidMultiblockMember = TargetRIP >= SymbolMinAddress && TargetRIP < SymbolMaxAddress;
// Forbid cross-page branches to both avoid massive (range-wise) code blocks in highly fragmented code and trying to decode unmapped branch targets
bool ValidMultiblockMember =
TargetRIP >= SymbolMinAddress && TargetRIP < std::min(FEXCore::AlignUp(InstEnd, FEXCore::Utils::FEX_PAGE_SIZE), SymbolMaxAddress);
#ifdef _M_ARM_64EC
ValidMultiblockMember = ValidMultiblockMember && !RtlIsEcCode(TargetRIP);
@@ -973,15 +977,10 @@ void Decoder::BranchTargetInMultiblockRange() {
MaxCondBranchBackwards = std::min(MaxCondBranchBackwards, TargetRIP);
// If we are conditional then a target can be the instruction past the conditional instruction
uint64_t FallthroughRIP = DecodeInst->PC + DecodeInst->InstSize;
if (!HasBlocks.contains(FallthroughRIP)) {
CurrentBlockTargets.insert(FallthroughRIP);
}
AddBranchTarget(InstEnd);
}
if (!HasBlocks.contains(TargetRIP)) {
CurrentBlockTargets.insert(TargetRIP);
}
AddBranchTarget(TargetRIP);
} else {
if (ExternalBranches) {
ExternalBranches->insert(TargetRIP);
@@ -989,19 +988,23 @@ void Decoder::BranchTargetInMultiblockRange() {
}
}
bool Decoder::BranchTargetCanContinue(bool FinalInstruction) const {
if (FinalInstruction) {
bool Decoder::InstCanContinue() const {
if (DecodeInst->PC + DecodeInst->InstSize == NextBlockStartAddress) {
return false;
}
if (!(DecodeInst->TableInfo->Flags & (FEXCore::X86Tables::InstFlags::FLAGS_BLOCK_END | FEXCore::X86Tables::InstFlags::FLAGS_SETS_RIP))) {
return true;
}
uint64_t TargetRIP = 0;
const uint8_t GPRSize = CTX->GetGPRSize();
const auto GPRSize = CTX->GetGPROpSize();
if (DecodeInst->OP == 0xE8) { // Call - immediate target
const uint64_t NextRIP = DecodeInst->PC + DecodeInst->InstSize;
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
if (GPRSize == 4) {
if (GPRSize == IR::OpSize::i32Bit) {
// If we are running a 32bit guest then wrap around addresses that go above 32bit
TargetRIP &= 0xFFFFFFFFU;
}
@@ -1017,6 +1020,59 @@ bool Decoder::BranchTargetCanContinue(bool FinalInstruction) const {
return false;
}
void Decoder::AddBranchTarget(uint64_t Target) {
if (VisitedBlocks.contains(Target)) {
return;
}
auto BlockSuccIt = std::lower_bound(BlockInfo.Blocks.begin(), BlockInfo.Blocks.end(), Target,
[](const auto& a, uint64_t Address) { return a.Entry < Address; });
LOGMAN_THROW_A_FMT(BlockSuccIt == BlockInfo.Blocks.end() || BlockSuccIt->Entry != Target, "unexpected");
if (BlockSuccIt != BlockInfo.Blocks.begin()) {
auto BlockIt = std::prev(BlockSuccIt);
if (BlockIt->Entry + BlockIt->Size > Target) {
uint64_t SplitIdx = 0;
uint64_t SplitAddr = BlockIt->Entry;
// Find the instruction boundary of the split
for (; SplitIdx < BlockIt->NumInstructions && SplitAddr < Target; SplitIdx++) {
SplitAddr += BlockIt->DecodedInstructions[SplitIdx].InstSize;
}
uint64_t SplitOffset = SplitAddr - BlockIt->Entry;
LOGMAN_THROW_A_FMT(SplitIdx != 0, "unexpected");
if (SplitAddr == Target) {
// Split at the boundary
DecodedBlocks SplitBlock {
.Entry = SplitAddr,
.Size = BlockIt->Size - SplitOffset,
.NumInstructions = BlockIt->NumInstructions - SplitIdx,
.DecodedInstructions = BlockIt->DecodedInstructions + SplitIdx,
.HasInvalidInstruction = BlockIt->HasInvalidInstruction,
};
BlockIt->Size = SplitOffset;
BlockIt->NumInstructions = SplitIdx;
BlockInfo.Blocks.insert(BlockSuccIt, SplitBlock);
} // else misaligned, leave as a branch out of the block
// If we split a block then the target has already been visited as part of that, if it was
// misaligned the jump will just leave the multiblock, mark it as visited to avoid running
// this code path again and just bail out early.
VisitedBlocks.insert(Target);
return;
}
}
CurrentBlockTargets.insert(Target);
if (Target >= DecodeInst->PC + DecodeInst->InstSize && Target < NextBlockStartAddress) {
NextBlockStartAddress = Target;
}
}
const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
constexpr uint64_t VSyscall_Base = 0xFFFF'FFFF'FF60'0000ULL;
constexpr uint64_t VSyscall_End = VSyscall_Base + 0x1000;
@@ -1040,7 +1096,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks.clear();
BlocksToDecode.clear();
HasBlocks.clear();
VisitedBlocks.clear();
// Reset internal state management
DecodedSize = 0;
MaxCondBranchForward = 0;
@@ -1078,30 +1134,61 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
}
bool EntryBlock {true};
bool FinalInstruction {false};
while (!BlocksToDecode.empty()) {
while (!FinalInstruction && !BlocksToDecode.empty()) {
auto BlockDecodeIt = BlocksToDecode.begin();
uint64_t RIPToDecode = *BlockDecodeIt;
BlockInfo.Blocks.emplace_back();
DecodedBlocks& CurrentBlockDecoding = BlockInfo.Blocks.back();
BlocksToDecode.erase(BlockDecodeIt);
VisitedBlocks.emplace(RIPToDecode);
CurrentBlockDecoding.Entry = RIPToDecode;
auto BlockSuccIt = std::lower_bound(BlockInfo.Blocks.begin(), BlockInfo.Blocks.end(), RIPToDecode,
[](const auto& a, uint64_t Address) { return a.Entry < Address; });
LOGMAN_THROW_A_FMT(BlockSuccIt == BlockInfo.Blocks.end() || BlockSuccIt->Entry != RIPToDecode, "unexpected");
NextBlockStartAddress = ~0ULL;
if (!BlocksToDecode.empty()) {
// We just erased the lowest, the front is then the second lowest
NextBlockStartAddress = *BlocksToDecode.begin();
}
if (BlockSuccIt != BlockInfo.Blocks.end() && BlockSuccIt->Entry < NextBlockStartAddress) {
NextBlockStartAddress = BlockSuccIt->Entry;
}
LOGMAN_THROW_A_FMT(NextBlockStartAddress > RIPToDecode, "unexpected");
// Insert the block now so it can be looked up and split if necessary on a backward edge
auto BlockIt = BlockInfo.Blocks.emplace(BlockSuccIt);
BlockIt->Entry = RIPToDecode;
BlockIt->Size = 0;
uint64_t PCOffset = 0;
uint64_t BlockNumberOfInstructions {};
uint64_t BlockStartOffset = DecodedSize;
bool EraseBlock = true; // Unset once the block contains an instruction
BlockIt->DecodedInstructions = &DecodedBuffer[BlockStartOffset];
BlockIt->NumInstructions = 0;
// Do a bit of pointer math to figure out where we are in code
InstStream = AdjustAddrForSpecialRegion(_InstStream, EntryPoint, RIPToDecode);
while (1) {
// MAX_INST_SIZE assumes worst case
auto OpMinAddress = RIPToDecode + PCOffset;
auto OpMaxAddress = OpMinAddress + MAX_INST_SIZE;
InstructionSize = 0;
auto OpMinPage = OpMinAddress & FEXCore::Utils::FEX_PAGE_MASK;
// MAX_INST_SIZE assumes worst case
auto OpAddress = RIPToDecode + PCOffset;
auto OpMaxAddress = OpAddress + MAX_INST_SIZE;
auto OpMinPage = OpAddress & FEXCore::Utils::FEX_PAGE_MASK;
auto OpMaxPage = OpMaxAddress & FEXCore::Utils::FEX_PAGE_MASK;
if (!EntryBlock && OpMinPage == OpMaxPage && PeekByte(0) == 0 && PeekByte(1) == 0) [[unlikely]] {
// End the multiblock early if we hit 2 consecutive null bytes (add [rax], al) in the same page with the
// assumption we are most likely trying to explore garbage code.
break;
}
if (OpMinPage != CurrentCodePage) {
CurrentCodePage = OpMinPage;
CodePages.insert(CurrentCodePage);
@@ -1112,64 +1199,66 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
CodePages.insert(CurrentCodePage);
}
bool ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
bool ErrorDuringDecoding = !DecodeInstruction(OpAddress);
uint64_t OpEndAddress = OpAddress + DecodeInst->InstSize;
if (ErrorDuringDecoding) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
CurrentBlockDecoding.HasInvalidInstruction = true;
BlockIt->HasInvalidInstruction = true;
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
}
if (!ErrorDuringDecoding) {
} else {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
auto TableInfo = DecodedBuffer[BlockStartOffset + BlockNumberOfInstructions].TableInfo;
auto TableInfo = DecodeInst->TableInfo;
if (!TableInfo || !TableInfo->OpcodeDispatcher) {
CurrentBlockDecoding.HasInvalidInstruction = true;
BlockIt->HasInvalidInstruction = true;
}
}
DecodedMinAddress = std::min(DecodedMinAddress, RIPToDecode + PCOffset);
DecodedMaxAddress = std::max(DecodedMaxAddress, RIPToDecode + PCOffset + DecodeInst->InstSize);
DecodedMinAddress = std::min(DecodedMinAddress, OpAddress);
DecodedMaxAddress = std::max(DecodedMaxAddress, OpEndAddress);
if (OpEndAddress > NextBlockStartAddress) {
// This instruction would overlap with another so skip adding it to the multiblock
break;
}
EraseBlock = false; // Block contains at least one valid instruction, so unset erase
++TotalInstructions;
++BlockNumberOfInstructions;
++DecodedSize;
++BlockIt->NumInstructions;
BlockIt->Size += DecodeInst->InstSize;
// Can not continue this block at all on invalid instruction
if (CurrentBlockDecoding.HasInvalidInstruction) [[unlikely]] {
if (BlockIt->HasInvalidInstruction) [[unlikely]] {
if (!EntryBlock) {
// In multiblock configurations, we can early terminate any non-entrypoint blocks with the expectation that this won't get hit.
// Improves compile-times.
// Just need to undo additions that this block decoding has caused.
TotalInstructions -= CurrentBlockDecoding.NumInstructions;
TotalInstructions -= BlockIt->NumInstructions;
DecodedSize = BlockStartOffset;
BlockNumberOfInstructions = 0;
InstStream -= PCOffset;
CurrentBlockTargets.clear();
EraseBlock = true;
}
break;
}
bool CanContinue = false;
if (!(DecodeInst->TableInfo->Flags & (FEXCore::X86Tables::InstFlags::FLAGS_BLOCK_END | FEXCore::X86Tables::InstFlags::FLAGS_SETS_RIP))) {
// If this isn't a block ender then we can keep going regardless
CanContinue = true;
// Check if we need to end the entire multiblock
FinalInstruction = DecodedSize >= MaxInst || DecodedSize >= DefaultDecodedBufferSize || TotalInstructions >= MaxInst;
if (FinalInstruction) {
break;
}
bool FinalInstruction = DecodedSize >= MaxInst || DecodedSize >= DefaultDecodedBufferSize || TotalInstructions >= MaxInst;
if (!InstCanContinue()) {
if (DecodeInst->TableInfo->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SETS_RIP) {
// If we have multiblock enabled
// If the branch target is within our multiblock range then we can keep going on
// We don't want to short circuit this since we want to calculate our ranges still
// NOTE: This will invalidate BlockIt, this is fine as we immediately break from the loop and EraseBlock cannot be true
BranchTargetInMultiblockRange();
}
if (DecodeInst->TableInfo->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SETS_RIP) {
// If we have multiblock enabled
// If the branch target is within our multiblock range then we can keep going on
// We don't want to short circuit this since we want to calculate our ranges still
BranchTargetInMultiblockRange();
// Bypass branches if we can continue through them in some cases.
CanContinue |= BranchTargetCanContinue(FinalInstruction);
}
if (FinalInstruction || !CanContinue) {
break;
}
@@ -1177,29 +1266,22 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
InstStream += DecodeInst->InstSize;
}
BlocksToDecode.merge(CurrentBlockTargets);
// NOTE: BlockIt is only valid here in the EraseBlock case
if (EraseBlock) {
BlockInfo.Blocks.erase(BlockIt);
} else {
BlocksToDecode.merge(CurrentBlockTargets);
}
CurrentBlockTargets.clear();
BlocksToDecode.erase(BlockDecodeIt);
HasBlocks.emplace(RIPToDecode);
// Copy over only the number of instructions we decoded
CurrentBlockDecoding.NumInstructions = BlockNumberOfInstructions;
CurrentBlockDecoding.DecodedInstructions = &DecodedBuffer[BlockStartOffset];
BlockInfo.TotalInstructionCount += BlockNumberOfInstructions;
EntryBlock = false;
}
BlockInfo.TotalInstructionCount = TotalInstructions;
for (auto CodePage : CodePages) {
AddContainedCodePage(PC, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
// sort for better branching
std::sort(BlockInfo.Blocks.begin(), BlockInfo.Blocks.end(),
[](const FEXCore::Frontend::Decoder::DecodedBlocks& a, const FEXCore::Frontend::Decoder::DecodedBlocks& b) {
return a.Entry < b.Entry;
});
}
} // namespace FEXCore::Frontend
+6 -2
View File
@@ -22,6 +22,7 @@ public:
// New Frontend decoding
struct DecodedBlocks final {
uint64_t Entry {};
uint64_t Size {};
uint64_t NumInstructions {};
FEXCore::X86Tables::DecodedInst* DecodedInstructions;
bool HasInvalidInstruction {};
@@ -70,7 +71,9 @@ private:
bool DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
bool BranchTargetCanContinue(bool FinalInstruction) const;
bool InstCanContinue() const;
void AddBranchTarget(uint64_t Target);
uint8_t ReadByte();
uint8_t PeekByte(uint8_t Offset) const;
@@ -102,11 +105,12 @@ private:
uint64_t SymbolMaxAddress {};
uint64_t SymbolMinAddress {~0ULL};
uint64_t SectionMaxAddress {~0ULL};
uint64_t NextBlockStartAddress {~0ULL};
DecodedBlockInformation BlockInfo;
fextl::set<uint64_t> CurrentBlockTargets;
fextl::set<uint64_t> BlocksToDecode;
fextl::set<uint64_t> HasBlocks;
fextl::set<uint64_t> VisitedBlocks;
fextl::set<uint64_t>* ExternalBranches {nullptr};
// ModRM rm decoding
@@ -5,18 +5,24 @@
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
#include "Interface/IR/IR.h"
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Profiler.h>
namespace FEXCore::CPU {
FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t FCW) {
FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t FCW, bool Force80BitPrecision = false) {
softfloat_state State {};
State.detectTininess = softfloat_tininess_afterRounding;
State.exceptionFlags = 0;
State.roundingPrecision = 80;
auto PC = (FCW >> 8) & 3;
switch (PC) {
case 0: State.roundingPrecision = 32; break;
case 2: State.roundingPrecision = 64; break;
case 3: State.roundingPrecision = 80; break;
case 1: LOGMAN_MSG_A_FMT("Invalid x87 precision mode, {}", PC);
if (!Force80BitPrecision) {
auto PC = (FCW >> 8) & 3;
switch (PC) {
case 0: State.roundingPrecision = 32; break;
case 2: State.roundingPrecision = 64; break;
case 3: State.roundingPrecision = 80; break;
case 1: LOGMAN_MSG_A_FMT("Invalid x87 precision mode, {}", PC);
}
}
auto RC = (FCW >> 10) & 3;
@@ -32,12 +38,14 @@ FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t FCW, float src) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle4(uint16_t FCW, float src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle8(uint16_t FCW, double src) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle8(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
}
@@ -45,7 +53,8 @@ struct OpHandlers<IR::OP_F80CVTTO> {
template<>
struct OpHandlers<IR::OP_F80CMP> {
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
bool eq, lt, nan;
@@ -67,37 +76,43 @@ struct OpHandlers<IR::OP_F80CMP> {
template<>
struct OpHandlers<IR::OP_F80CVT> {
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToF32(&State);
return X80SoftFloat(src).ToF32(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToF64(&State);
return X80SoftFloat(src).ToF64(&State);
}
};
template<>
struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI16(&State);
return X80SoftFloat(src).ToI16(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI32(&State);
return X80SoftFloat(src).ToI32(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI64(&State);
return X80SoftFloat(src).ToI64(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
auto rv = extF80_to_i32(&State, src, softfloat_round_minMag, false);
auto rv = extF80_to_i32(&State, X80SoftFloat(src), softfloat_round_minMag, false);
if (rv > INT16_MAX || rv < INT16_MIN) {
///< Indefinite value for 16-bit conversions.
@@ -107,55 +122,63 @@ struct OpHandlers<IR::OP_F80CVTINT> {
}
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i32(&State, src, softfloat_round_minMag, false);
return extF80_to_i32(&State, X80SoftFloat(src), softfloat_round_minMag, false);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t FCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i64(&State, src, softfloat_round_minMag, false);
return extF80_to_i64(&State, X80SoftFloat(src), softfloat_round_minMag, false);
}
};
template<>
struct OpHandlers<IR::OP_F80CVTTOINT> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle2(uint16_t FCW, int16_t src) {
return src;
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle2(uint16_t FCW, int16_t src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return X80SoftFloat(src);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t FCW, int32_t src) {
return src;
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle4(uint16_t FCW, int32_t src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return X80SoftFloat(src);
}
};
template<>
struct OpHandlers<IR::OP_F80ROUND> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FRNDINT(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::F2XM1(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80TAN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FTAN(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SQRT> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSQRT(&State, Src1);
}
@@ -163,37 +186,42 @@ struct OpHandlers<IR::OP_F80SQRT> {
template<>
struct OpHandlers<IR::OP_F80SIN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FSIN(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80COS> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FCOS(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_EXP> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return X80SoftFloat::FXTRACT_EXP(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_SIG> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return X80SoftFloat::FXTRACT_SIG(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80ADD> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FADD(&State, Src1, Src2);
}
@@ -201,7 +229,8 @@ struct OpHandlers<IR::OP_F80ADD> {
template<>
struct OpHandlers<IR::OP_F80SUB> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSUB(&State, Src1, Src2);
}
@@ -209,7 +238,8 @@ struct OpHandlers<IR::OP_F80SUB> {
template<>
struct OpHandlers<IR::OP_F80MUL> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FMUL(&State, Src1, Src2);
}
@@ -217,7 +247,8 @@ struct OpHandlers<IR::OP_F80MUL> {
template<>
struct OpHandlers<IR::OP_F80DIV> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FDIV(&State, Src1, Src2);
}
@@ -225,103 +256,117 @@ struct OpHandlers<IR::OP_F80DIV> {
template<>
struct OpHandlers<IR::OP_F80FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FYL2X(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FATAN(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FREM1(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FREM(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FSCALE(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F64SIN> {
static double handle(uint16_t FCW, double src) {
static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return sin(src);
}
};
template<>
struct OpHandlers<IR::OP_F64COS> {
static double handle(uint16_t FCW, double src) {
static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return cos(src);
}
};
template<>
struct OpHandlers<IR::OP_F64TAN> {
static double handle(uint16_t FCW, double src) {
static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return tan(src);
}
};
template<>
struct OpHandlers<IR::OP_F64F2XM1> {
static double handle(uint16_t FCW, double src) {
static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return exp2(src) - 1.0;
}
};
template<>
struct OpHandlers<IR::OP_F64ATAN> {
static double handle(uint16_t FCW, double src1, double src2) {
static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return atan2(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM> {
static double handle(uint16_t FCW, double src1, double src2) {
static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return fmod(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM1> {
static double handle(uint16_t FCW, double src1, double src2) {
static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return remainder(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2X> {
static double handle(uint16_t FCW, double src1, double src2) {
static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return src2 * log2(src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
static double handle(uint16_t FCW, double src1, double src2) {
static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
if (src1 == 0.0) { // src1 might be +/- zero
return src1; // this will return negative or positive zero if when appropriate
}
@@ -332,7 +377,9 @@ struct OpHandlers<IR::OP_F64SCALE> {
template<>
struct OpHandlers<IR::OP_F80BCDSTORE> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1q, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
X80SoftFloat Src1 = Src1q;
softfloat_state State = SoftFloatStateFromFCW(FCW);
bool Negative = Src1.Sign;
@@ -373,7 +420,8 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
template<>
struct OpHandlers<IR::OP_F80BCDLOAD> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src) {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
uint8_t* Src1 = reinterpret_cast<uint8_t*>(&Src);
uint64_t BCD {};
// We walk through each uint8_t and pull out the BCD encoding
@@ -1,3 +1,4 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Core/CoreState.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
@@ -15,13 +16,14 @@ static FallbackInfo GetFallbackInfo(R (*fn)(Args...), FEXCore::Core::FallbackHan
}
template<>
FallbackInfo GetFallbackInfo(double (*fn)(uint16_t, double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_I16_F64, (void*)fn, HandlerIndex, false};
FallbackInfo GetFallbackInfo(double (*fn)(uint16_t, double, FEXCore::Core::CpuStateFrame*), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_I16_F64_PTR, (void*)fn, HandlerIndex, false};
}
template<>
FallbackInfo GetFallbackInfo(double (*fn)(uint16_t, double, double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_I16_F64_F64, (void*)fn, HandlerIndex, false};
FallbackInfo
GetFallbackInfo(double (*fn)(uint16_t, double, double, FEXCore::Core::CpuStateFrame*), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_I16_F64_F64_PTR, (void*)fn, HandlerIndex, false};
}
void InterpreterOps::FillFallbackIndexPointers(uint64_t* Info) {
@@ -79,18 +81,18 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t* Info) {
}
bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::IROp_Header* IROp, FallbackInfo* Info) {
uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
switch (IROp->Op) {
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->SrcSize) {
case 4: {
*Info = {FABI_F80_I16_F32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4, SupportsPreserveAllABI};
case IR::OpSize::i32Bit: {
*Info = {FABI_F80_I16_F32_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F80_I16_F64, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8, SupportsPreserveAllABI};
case IR::OpSize::i64Bit: {
*Info = {FABI_F80_I16_F64_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
@@ -99,12 +101,12 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = {FABI_F32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4, SupportsPreserveAllABI};
case IR::OpSize::i32Bit: {
*Info = {FABI_F32_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8, SupportsPreserveAllABI};
case IR::OpSize::i64Bit: {
*Info = {FABI_F64_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
@@ -115,30 +117,33 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
case IR::OpSize::i16Bit: {
if (Op->Truncate) {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2,
*Info = {FABI_I16_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2,
SupportsPreserveAllABI};
} else {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2, SupportsPreserveAllABI};
*Info = {FABI_I16_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2,
SupportsPreserveAllABI};
}
return true;
}
case 4: {
case IR::OpSize::i32Bit: {
if (Op->Truncate) {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4,
*Info = {FABI_I32_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4,
SupportsPreserveAllABI};
} else {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4, SupportsPreserveAllABI};
*Info = {FABI_I32_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4,
SupportsPreserveAllABI};
}
return true;
}
case 8: {
case IR::OpSize::i64Bit: {
if (Op->Truncate) {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8,
*Info = {FABI_I64_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8,
SupportsPreserveAllABI};
} else {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8, SupportsPreserveAllABI};
*Info = {FABI_I64_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8,
SupportsPreserveAllABI};
}
return true;
}
@@ -147,7 +152,7 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
break;
}
case IR::OP_F80CMP: {
*Info = {FABI_I64_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle,
*Info = {FABI_I64_I16_F80_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle,
(Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP), SupportsPreserveAllABI};
return true;
}
@@ -156,12 +161,14 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->SrcSize) {
case 2: {
*Info = {FABI_F80_I16_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2, SupportsPreserveAllABI};
case IR::OpSize::i16Bit: {
*Info = {FABI_F80_I16_I16_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2,
SupportsPreserveAllABI};
return true;
}
case 4: {
*Info = {FABI_F80_I16_I32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4, SupportsPreserveAllABI};
case IR::OpSize::i32Bit: {
*Info = {FABI_F80_I16_I32_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4,
SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
@@ -169,16 +176,16 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
break;
}
#define COMMON_UNARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
#define COMMON_UNARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_F80_PTR, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
#define COMMON_F64_OP(OP) \
@@ -229,7 +236,7 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
SupportsPreserveAllABI};
return true;
case IR::OP_VPCMPISTRX:
*Info = {FABI_I32_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX, SupportsPreserveAllABI};
*Info = {FABI_I32_V128_V128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX, SupportsPreserveAllABI};
return true;
default: break;
@@ -0,0 +1,91 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/Interpreter/Fallbacks/VectorFallbacks.h"
#include "Interface/IR/IR.h"
#ifdef _M_ARM_64
#include <arm_neon.h>
#endif
#include <cstring>
namespace FEXCore::CPU {
#ifdef _M_ARM_64
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetImplicitLength(FEXCore::VectorRegType data, uint16_t control) {
const auto is_using_words = (control & 1) != 0;
if (is_using_words) {
uint16x8_t a = vreinterpretq_u16_u8(data);
uint16x8_t VIndexes {};
const uint16x8_t VIndex16 = vdupq_n_u16(8);
uint16_t Indexes[8] = {
0, 1, 2, 3, 4, 5, 6, 7,
};
memcpy(&VIndexes, Indexes, sizeof(VIndexes));
auto MaskResult = vceqzq_u16(a);
auto SelectResult = vbslq_u16(MaskResult, VIndexes, VIndex16);
return vminvq_u16(SelectResult);
} else {
uint8x16_t VIndexes {};
const uint8x16_t VIndex16 = vdupq_n_u8(16);
uint8_t Indexes[16] = {
0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15,
};
memcpy(&VIndexes, Indexes, sizeof(VIndexes));
auto MaskResult = vceqzq_u8(data);
auto SelectResult = vbslq_u8(MaskResult, VIndexes, VIndex16);
return vminvq_u8(SelectResult);
}
}
#else
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetImplicitLength(FEXCore::VectorRegType data, uint16_t control) {
const auto* data_u8 = reinterpret_cast<const uint8_t*>(&data);
const auto is_using_words = (control & 1) != 0;
int32_t length = 0;
if (is_using_words) {
const auto get_word = [data_u8](int32_t index) {
const auto* src = data_u8 + (index * sizeof(uint16_t));
uint16_t element {};
std::memcpy(&element, src, sizeof(uint16_t));
return element;
};
while (length < 8 && get_word(length) != 0) {
length++;
}
} else {
while (length < 16 && data_u8[length] != 0) {
length++;
}
}
return length;
}
#endif
// Essentially the same in terms of behavior with VPCMPESTRX instructions,
// with the only difference being that the length of the string is encoded
// as part of the data vectors passed in.
//
// i.e. Length is determined by the presence of a NUL (all-zero) character
// within the data.
//
// If no NUL character exists, then the length of the strings are assumed
// to be the max length possible for the given character size specified
// in the control flags (16 characters for 8-bit, and 8 characters for 16-bit).
//
FEXCORE_PRESERVE_ALL_ATTR uint32_t OpHandlers<IR::OP_VPCMPISTRX>::handle(FEXCore::VectorRegType lhs, FEXCore::VectorRegType rhs, uint16_t control) {
// Subtract by 1 in order to make validity limits 0-based
const auto valid_lhs = GetImplicitLength(lhs, control) - 1;
const auto valid_rhs = GetImplicitLength(rhs, control) - 1;
__uint128_t lhs_i;
memcpy(&lhs_i, &lhs, sizeof(lhs_i));
__uint128_t rhs_i;
memcpy(&rhs_i, &rhs, sizeof(rhs_i));
return OpHandlers<IR::OP_VPCMPESTRX>::MainBody(lhs_i, valid_lhs, rhs_i, valid_rhs, control);
}
} // namespace FEXCore::CPU
@@ -6,9 +6,9 @@
#include <cstdlib>
#include <cstring>
#include <FEXCore/IR/IR.h>
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
#include "Interface/IR/IR.h"
#include "Common/VectorRegType.h"
namespace FEXCore::CPU {
@@ -344,51 +344,7 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
template<>
struct OpHandlers<IR::OP_VPCMPISTRX> {
// Essentially the same in terms of behavior with VPCMPESTRX instructions,
// with the only difference being that the length of the string is encoded
// as part of the data vectors passed in.
//
// i.e. Length is determined by the presence of a NUL (all-zero) character
// within the data.
//
// If no NUL character exists, then the length of the strings are assumed
// to be the max length possible for the given character size specified
// in the control flags (16 characters for 8-bit, and 8 characters for 16-bit).
//
FEXCORE_PRESERVE_ALL_ATTR static uint32_t handle(__uint128_t lhs, __uint128_t rhs, uint16_t control) {
// Subtract by 1 in order to make validity limits 0-based
const auto valid_lhs = GetImplicitLength(lhs, control) - 1;
const auto valid_rhs = GetImplicitLength(rhs, control) - 1;
return OpHandlers<IR::OP_VPCMPESTRX>::MainBody(lhs, valid_lhs, rhs, valid_rhs, control);
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetImplicitLength(const __uint128_t& data, uint16_t control) {
const auto* data_u8 = reinterpret_cast<const uint8_t*>(&data);
const auto is_using_words = (control & 1) != 0;
int32_t length = 0;
if (is_using_words) {
const auto get_word = [data_u8](int32_t index) {
const auto* src = data_u8 + (index * sizeof(uint16_t));
uint16_t element {};
std::memcpy(&element, src, sizeof(uint16_t));
return element;
};
while (length < 8 && get_word(length) != 0) {
length++;
}
} else {
while (length < 16 && data_u8[length] != 0) {
length++;
}
}
return length;
}
FEXCORE_PRESERVE_ALL_ATTR static uint32_t handle(VectorRegType lhs, VectorRegType rhs, uint16_t control);
};
} // namespace FEXCore::CPU
@@ -1,8 +1,6 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <array>
#include <cstddef>
#include <cstdint>
#include <FEXCore/Core/CoreState.h>
@@ -16,22 +14,22 @@ struct IROp_Header;
namespace FEXCore::CPU {
enum FallbackABI {
FABI_UNKNOWN,
FABI_F80_I16_F32,
FABI_F80_I16_F64,
FABI_F80_I16_I16,
FABI_F80_I16_I32,
FABI_F32_I16_F80,
FABI_F64_I16_F80,
FABI_F64_I16_F64,
FABI_F64_I16_F64_F64,
FABI_I16_I16_F80,
FABI_I32_I16_F80,
FABI_I64_I16_F80,
FABI_I64_I16_F80_F80,
FABI_F80_I16_F80,
FABI_F80_I16_F80_F80,
FABI_F80_I16_F32_PTR,
FABI_F80_I16_F64_PTR,
FABI_F80_I16_I16_PTR,
FABI_F80_I16_I32_PTR,
FABI_F32_I16_F80_PTR,
FABI_F64_I16_F80_PTR,
FABI_F64_I16_F64_PTR,
FABI_F64_I16_F64_F64_PTR,
FABI_I16_I16_F80_PTR,
FABI_I32_I16_F80_PTR,
FABI_I64_I16_F80_PTR,
FABI_I64_I16_F80_F80_PTR,
FABI_F80_I16_F80_PTR,
FABI_F80_I16_F80_F80_PTR,
FABI_I32_I64_I64_I128_I128_I16,
FABI_I32_I128_I128_I16,
FABI_I32_V128_V128_I16,
};
struct FallbackInfo {
@@ -8,7 +8,7 @@ $end_info$
#include "CodeEmitter/Emitter.h"
#include "FEXCore/IR/IR.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
@@ -54,8 +54,8 @@ DEF_OP(EntrypointOffset) {
auto Constant = Entry + Op->Offset;
auto Dst = GetReg(Node);
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
const auto OpSize = IROp->Size;
if (OpSize == IR::OpSize::i32Bit) {
Mask = 0xFFFF'FFFFULL;
}
@@ -92,10 +92,10 @@ DEF_OP(AddNZCV) {
uint64_t Const;
if (IsInlineConstant(Op->Src2, &Const)) {
LOGMAN_THROW_AA_FMT(IROp->Size >= 4, "Constant not allowed here");
LOGMAN_THROW_A_FMT(IROp->Size >= IR::OpSize::i32Bit, "Constant not allowed here");
cmn(EmitSize, Src1, Const);
} else if (IROp->Size < 4) {
unsigned Shift = 32 - (8 * IROp->Size);
} else if (IROp->Size < IR::OpSize::i32Bit) {
unsigned Shift = 32 - IR::OpSizeAsBits(IROp->Size);
lsl(ARMEmitter::Size::i32Bit, TMP1, Src1, Shift);
cmn(EmitSize, TMP1, GetReg(Op->Src2.ID()), ARMEmitter::ShiftType::LSL, Shift);
@@ -165,7 +165,7 @@ DEF_OP(TestNZ) {
// Shift the sign bit into place, clearing out the garbage in upper bits.
// Adding zero does an effective test, setting NZ according to the result and
// zeroing CV.
if (IROp->Size < 4) {
if (IROp->Size < IR::OpSize::i32Bit) {
// Cheaper to and+cmn than to lsl+lsl+tst, so do the and ourselves if
// needed.
if (Op->Src1 != Op->Src2) {
@@ -179,7 +179,7 @@ DEF_OP(TestNZ) {
Src1 = TMP1;
}
unsigned Shift = 32 - (IROp->Size * 8);
unsigned Shift = 32 - IR::OpSizeAsBits(IROp->Size);
cmn(EmitSize, ARMEmitter::Reg::zr, Src1, ARMEmitter::ShiftType::LSL, Shift);
} else {
if (IsInlineConstant(Op->Src2, &Const)) {
@@ -193,16 +193,16 @@ DEF_OP(TestNZ) {
DEF_OP(TestZ) {
auto Op = IROp->C<IR::IROp_TestZ>();
LOGMAN_THROW_AA_FMT(IROp->Size < 4, "TestNZ used at higher sizes");
LOGMAN_THROW_A_FMT(IROp->Size < IR::OpSize::i32Bit, "TestNZ used at higher sizes");
const auto EmitSize = ARMEmitter::Size::i32Bit;
uint64_t Const;
uint64_t Mask = IROp->Size == 8 ? ~0ULL : ((1ull << (IROp->Size * 8)) - 1);
uint64_t Mask = IROp->Size == IR::OpSize::i64Bit ? ~0ULL : ((1ull << IR::OpSizeAsBits(IROp->Size)) - 1);
auto Src1 = GetReg(Op->Src1.ID());
if (IsInlineConstant(Op->Src2, &Const)) {
// We can promote 8/16-bit tests to 32-bit since the constant is masked.
LOGMAN_THROW_AA_FMT(!(Const & ~Mask), "constant is already masked");
LOGMAN_THROW_A_FMT(!(Const & ~Mask), "constant is already masked");
tst(EmitSize, Src1, Const);
} else {
const auto Src2 = GetReg(Op->Src2.ID());
@@ -223,25 +223,25 @@ DEF_OP(SubShift) {
DEF_OP(SubNZCV) {
auto Op = IROp->C<IR::IROp_SubNZCV>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
uint64_t Const;
if (IsInlineConstant(Op->Src2, &Const)) {
LOGMAN_THROW_AA_FMT(OpSize >= 4, "Constant not allowed here");
LOGMAN_THROW_A_FMT(OpSize >= IR::OpSize::i32Bit, "Constant not allowed here");
cmp(EmitSize, GetReg(Op->Src1.ID()), Const);
} else {
unsigned Shift = OpSize < 4 ? (32 - (8 * OpSize)) : 0;
unsigned Shift = OpSize < IR::OpSize::i32Bit ? (32 - IR::OpSizeAsBits(OpSize)) : 0;
ARMEmitter::Register ShiftedSrc1 = GetZeroableReg(Op->Src1);
// Shift to fix flags for <32-bit ops.
// Any shift of zero is still zero so optimize out silly zero shifts.
if (OpSize < 4 && ShiftedSrc1 != ARMEmitter::Reg::zr) {
if (OpSize < IR::OpSize::i32Bit && ShiftedSrc1 != ARMEmitter::Reg::zr) {
lsl(ARMEmitter::Size::i32Bit, TMP1, ShiftedSrc1, Shift);
ShiftedSrc1 = TMP1;
}
if (OpSize < 4) {
if (OpSize < IR::OpSize::i32Bit) {
cmp(EmitSize, ShiftedSrc1, GetReg(Op->Src2.ID()), ARMEmitter::ShiftType::LSL, Shift);
} else {
cmp(EmitSize, ShiftedSrc1, GetReg(Op->Src2.ID()));
@@ -286,10 +286,10 @@ DEF_OP(SetSmallNZV) {
auto Op = IROp->C<IR::IROp_SetSmallNZV>();
LOGMAN_THROW_A_FMT(CTX->HostFeatures.SupportsFlagM, "Unsupported flagm op");
const uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 1 || OpSize == 2, "Unsupported {} size: {}", __func__, OpSize);
const auto OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i8Bit || OpSize == IR::OpSize::i16Bit, "Unsupported {} size: {}", __func__, OpSize);
if (OpSize == 1) {
if (OpSize == IR::OpSize::i8Bit) {
setf8(GetReg(Op->Src.ID()).W());
} else {
setf16(GetReg(Op->Src.ID()).W());
@@ -401,20 +401,20 @@ DEF_OP(Div) {
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
auto Src1 = GetReg(Op->Src1.ID());
auto Src2 = GetReg(Op->Src2.ID());
if (OpSize == 1) {
if (OpSize == IR::OpSize::i8Bit) {
sxtb(EmitSize, TMP1, Src1);
sxtb(EmitSize, TMP2, Src2);
Src1 = TMP1;
Src2 = TMP2;
} else if (OpSize == 2) {
} else if (OpSize == IR::OpSize::i16Bit) {
sxth(EmitSize, TMP1, Src1);
sxth(EmitSize, TMP2, Src2);
@@ -430,20 +430,20 @@ DEF_OP(UDiv) {
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
auto Src1 = GetReg(Op->Src1.ID());
auto Src2 = GetReg(Op->Src2.ID());
if (OpSize == 1) {
if (OpSize == IR::OpSize::i8Bit) {
uxtb(EmitSize, TMP1, Src1);
uxtb(EmitSize, TMP2, Src2);
Src1 = TMP1;
Src2 = TMP2;
} else if (OpSize == 2) {
} else if (OpSize == IR::OpSize::i16Bit) {
uxth(EmitSize, TMP1, Src1);
uxth(EmitSize, TMP2, Src2);
@@ -458,20 +458,20 @@ DEF_OP(Rem) {
auto Op = IROp->C<IR::IROp_Rem>();
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
auto Src1 = GetReg(Op->Src1.ID());
auto Src2 = GetReg(Op->Src2.ID());
if (OpSize == 1) {
if (OpSize == IR::OpSize::i8Bit) {
sxtb(EmitSize, TMP1, Src1);
sxtb(EmitSize, TMP2, Src2);
Src1 = TMP1;
Src2 = TMP2;
} else if (OpSize == 2) {
} else if (OpSize == IR::OpSize::i16Bit) {
sxth(EmitSize, TMP1, Src1);
sxth(EmitSize, TMP2, Src2);
@@ -487,20 +487,20 @@ DEF_OP(URem) {
auto Op = IROp->C<IR::IROp_URem>();
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
auto Src1 = GetReg(Op->Src1.ID());
auto Src2 = GetReg(Op->Src2.ID());
if (OpSize == 1) {
if (OpSize == IR::OpSize::i8Bit) {
uxtb(EmitSize, TMP1, Src1);
uxtb(EmitSize, TMP2, Src2);
Src1 = TMP1;
Src2 = TMP2;
} else if (OpSize == 2) {
} else if (OpSize == IR::OpSize::i16Bit) {
uxth(EmitSize, TMP1, Src1);
uxth(EmitSize, TMP2, Src2);
@@ -514,15 +514,15 @@ DEF_OP(URem) {
DEF_OP(MulH) {
auto Op = IROp->C<IR::IROp_MulH>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto Dst = GetReg(Node);
const auto Src1 = GetReg(Op->Src1.ID());
const auto Src2 = GetReg(Op->Src2.ID());
if (OpSize == 4) {
if (OpSize == IR::OpSize::i32Bit) {
sxtw(TMP1, Src1.W());
sxtw(TMP2, Src2.W());
mul(ARMEmitter::Size::i32Bit, Dst, TMP1, TMP2);
@@ -534,15 +534,15 @@ DEF_OP(MulH) {
DEF_OP(UMulH) {
auto Op = IROp->C<IR::IROp_UMulH>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto Dst = GetReg(Node);
const auto Src1 = GetReg(Op->Src1.ID());
const auto Src2 = GetReg(Op->Src2.ID());
if (OpSize == 4) {
if (OpSize == IR::OpSize::i32Bit) {
uxtw(ARMEmitter::Size::i64Bit, TMP1, Src1);
uxtw(ARMEmitter::Size::i64Bit, TMP2, Src2);
mul(ARMEmitter::Size::i64Bit, Dst, TMP1, TMP2);
@@ -593,7 +593,7 @@ DEF_OP(Ornror) {
DEF_OP(AndWithFlags) {
auto Op = IROp->C<IR::IROp_AndWithFlags>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
uint64_t Const;
@@ -601,7 +601,7 @@ DEF_OP(AndWithFlags) {
auto Src1 = GetReg(Op->Src1.ID());
// See TestNZ
if (OpSize < 4) {
if (OpSize < IR::OpSize::i32Bit) {
if (IsInlineConstant(Op->Src2, &Const)) {
and_(EmitSize, Dst, Src1, Const);
} else {
@@ -614,7 +614,7 @@ DEF_OP(AndWithFlags) {
}
}
unsigned Shift = 32 - (OpSize * 8);
unsigned Shift = 32 - IR::OpSizeAsBits(OpSize);
cmn(EmitSize, ARMEmitter::Reg::zr, Dst, ARMEmitter::ShiftType::LSL, Shift);
} else {
if (IsInlineConstant(Op->Src2, &Const)) {
@@ -640,7 +640,7 @@ DEF_OP(XornShift) {
DEF_OP(Ashr) {
auto Op = IROp->C<IR::IROp_Ashr>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
@@ -648,29 +648,29 @@ DEF_OP(Ashr) {
uint64_t Const;
if (IsInlineConstant(Op->Src2, &Const)) {
if (OpSize >= 4) {
if (OpSize >= IR::OpSize::i32Bit) {
asr(EmitSize, Dst, Src1, (unsigned int)Const);
} else {
sbfx(EmitSize, TMP1, Src1, 0, OpSize * 8);
sbfx(EmitSize, TMP1, Src1, 0, IR::OpSizeAsBits(OpSize));
asr(EmitSize, Dst, TMP1, (unsigned int)Const);
ubfx(EmitSize, Dst, Dst, 0, OpSize * 8);
ubfx(EmitSize, Dst, Dst, 0, IR::OpSizeAsBits(OpSize));
}
} else {
const auto Src2 = GetReg(Op->Src2.ID());
if (OpSize >= 4) {
if (OpSize >= IR::OpSize::i32Bit) {
asrv(EmitSize, Dst, Src1, Src2);
} else {
sbfx(EmitSize, TMP1, Src1, 0, OpSize * 8);
sbfx(EmitSize, TMP1, Src1, 0, IR::OpSizeAsBits(OpSize));
asrv(EmitSize, Dst, TMP1, Src2);
ubfx(EmitSize, Dst, Dst, 0, OpSize * 8);
ubfx(EmitSize, Dst, Dst, 0, IR::OpSizeAsBits(OpSize));
}
}
}
DEF_OP(ShiftFlags) {
auto Op = IROp->C<IR::IROp_ShiftFlags>();
const uint8_t OpSize = Op->Size;
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto OpSize = Op->Size;
const auto EmitSize = OpSize == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto PFOutput = GetReg(Node);
const auto PFInput = GetReg(Op->PFInput.ID());
@@ -690,16 +690,16 @@ DEF_OP(ShiftFlags) {
// We need to mask the source before comparing it. We don't just skip flag
// updates for Src2=0 but anything that masks to zero.
and_(ARMEmitter::Size::i32Bit, TMP1, Src2, OpSize == 8 ? 0x3f : 0x1f);
and_(ARMEmitter::Size::i32Bit, TMP1, Src2, OpSize == IR::OpSize::i64Bit ? 0x3f : 0x1f);
ARMEmitter::SingleUseForwardLabel Done;
ARMEmitter::ForwardLabel Done;
cbz(EmitSize, TMP1, &Done);
{
// PF/SF/ZF/OF
if (OpSize >= 4) {
if (OpSize >= IR::OpSize::i32Bit) {
ands(EmitSize, PFTemp, Dst, Dst);
} else {
unsigned Shift = 32 - (OpSize * 8);
unsigned Shift = 32 - (IR::OpSizeToSize(OpSize) * 8);
cmn(EmitSize, ARMEmitter::Reg::zr, Dst, ARMEmitter::ShiftType::LSL, Shift);
mov(ARMEmitter::Size::i64Bit, PFTemp, Dst);
}
@@ -709,12 +709,12 @@ DEF_OP(ShiftFlags) {
// Extract the last bit shifted in to CF
if (Op->Shift == IR::ShiftType::LSL) {
if (OpSize >= 4) {
if (OpSize >= IR::OpSize::i32Bit) {
neg(EmitSize, CFWord, Src2);
lsrv(EmitSize, CFWord, Src1, CFWord);
} else {
CFWord = Dst.X();
CFBit = (OpSize * 8);
CFBit = IR::OpSizeToSize(OpSize) * 8;
}
} else {
sub(ARMEmitter::Size::i64Bit, CFWord, Src2, 1);
@@ -737,7 +737,7 @@ DEF_OP(ShiftFlags) {
rmif(CFWord, (CFBit - 1) % 64, (1 << 1) /* C */);
if (SetOF) {
rmif(TMP3, OpSize * 8 - 1, (1 << 0) /* V */);
rmif(TMP3, IR::OpSizeToSize(OpSize) * 8 - 1, (1 << 0) /* V */);
}
} else {
mrs(TMP2, ARMEmitter::SystemRegister::NZCV);
@@ -750,7 +750,7 @@ DEF_OP(ShiftFlags) {
bfi(ARMEmitter::Size::i32Bit, TMP2, CFWord, 29 /* C */, 1);
if (SetOF) {
lsr(EmitSize, TMP3, TMP3, OpSize * 8 - 1);
lsr(EmitSize, TMP3, TMP3, IR::OpSizeToSize(OpSize) * 8 - 1);
bfi(ARMEmitter::Size::i32Bit, TMP2, TMP3, 28 /* V */, 1);
}
@@ -770,14 +770,14 @@ DEF_OP(RotateFlags) {
const auto Result = GetReg(Op->Result.ID());
const auto Shift = GetReg(Op->Shift.ID());
const bool Left = Op->Left;
const auto EmitSize = Op->Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto EmitSize = Op->Size == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
// If shift=0, flags are unaffected. Wrap the whole implementation in a cbz.
ARMEmitter::SingleUseForwardLabel Done;
ARMEmitter::ForwardLabel Done;
cbz(EmitSize, Shift, &Done);
{
// Extract the last bit shifted in to CF
const auto BitSize = Op->Size * 8;
const auto BitSize = IR::OpSizeToSize(Op->Size) * 8;
unsigned CFBit = Left ? 0 : BitSize - 1;
// For ROR, OF is the XOR of the new CF bit and the most significant bit of the result.
@@ -862,7 +862,7 @@ DEF_OP(PDep) {
const auto T1 = TMP4.R();
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::SingleUseForwardLabel Done;
ARMEmitter::ForwardLabel Done;
// First, copy the input/mask, since we'll be clobbering. Copy as 64-bit to
// make this 0-uop on Firestorm.
@@ -897,7 +897,7 @@ DEF_OP(PDep) {
DEF_OP(PExt) {
auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const auto OpSizeBitsM1 = (OpSize * 8) - 1;
const auto OpSizeBitsM1 = IR::OpSizeAsBits(OpSize) - 1;
const auto EmitSize = ConvertSize48(IROp);
const auto Input = GetReg(Op->Input.ID());
@@ -922,9 +922,9 @@ DEF_OP(PExt) {
const auto BitReg = TMP2;
const auto ValueReg = TMP3;
ARMEmitter::SingleUseForwardLabel EarlyExit;
ARMEmitter::ForwardLabel EarlyExit;
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::SingleUseForwardLabel Done;
ARMEmitter::ForwardLabel Done;
cbz(EmitSize, Mask, &EarlyExit);
mov(EmitSize, MaskReg, Mask);
@@ -952,8 +952,8 @@ DEF_OP(PExt) {
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
const uint8_t OpSize = IROp->Size;
const auto EmitSize = OpSize >= 4 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto OpSize = IROp->Size;
const auto EmitSize = OpSize >= IR::OpSize::i32Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Dst = GetReg(Node);
const auto Upper = GetReg(Op->Upper.ID());
@@ -963,14 +963,14 @@ DEF_OP(LDiv) {
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
switch (OpSize) {
case 2: {
case IR::OpSize::i16Bit: {
uxth(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 16, 16);
sxth(EmitSize, TMP2, Divisor);
sdiv(EmitSize, Dst, TMP1, TMP2);
break;
}
case 4: {
case IR::OpSize::i32Bit: {
// TODO: 32-bit operation should be guaranteed not to leave garbage in the upper bits.
mov(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 32, 32);
@@ -978,9 +978,9 @@ DEF_OP(LDiv) {
sdiv(EmitSize, Dst, TMP1, TMP2);
break;
}
case 8: {
ARMEmitter::SingleUseForwardLabel Only64Bit {};
ARMEmitter::SingleUseForwardLabel LongDIVRet {};
case IR::OpSize::i64Bit: {
ARMEmitter::ForwardLabel Only64Bit {};
ARMEmitter::ForwardLabel LongDIVRet {};
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
@@ -1022,8 +1022,8 @@ DEF_OP(LDiv) {
DEF_OP(LUDiv) {
auto Op = IROp->C<IR::IROp_LUDiv>();
const uint8_t OpSize = IROp->Size;
const auto EmitSize = OpSize >= 4 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto OpSize = IROp->Size;
const auto EmitSize = OpSize >= IR::OpSize::i32Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Dst = GetReg(Node);
const auto Upper = GetReg(Op->Upper.ID());
@@ -1033,22 +1033,22 @@ DEF_OP(LUDiv) {
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64=
switch (OpSize) {
case 2: {
case IR::OpSize::i16Bit: {
uxth(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 16, 16);
udiv(EmitSize, Dst, TMP1, Divisor);
break;
}
case 4: {
case IR::OpSize::i32Bit: {
// TODO: 32-bit operation should be guaranteed not to leave garbage in the upper bits.
mov(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 32, 32);
udiv(EmitSize, Dst, TMP1, Divisor);
break;
}
case 8: {
ARMEmitter::SingleUseForwardLabel Only64Bit {};
ARMEmitter::SingleUseForwardLabel LongDIVRet {};
case IR::OpSize::i64Bit: {
ARMEmitter::ForwardLabel Only64Bit {};
ARMEmitter::ForwardLabel LongDIVRet {};
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
@@ -1086,8 +1086,8 @@ DEF_OP(LUDiv) {
DEF_OP(LRem) {
auto Op = IROp->C<IR::IROp_LRem>();
const uint8_t OpSize = IROp->Size;
const auto EmitSize = OpSize >= 4 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto OpSize = IROp->Size;
const auto EmitSize = OpSize >= IR::OpSize::i32Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Dst = GetReg(Node);
const auto Upper = GetReg(Op->Upper.ID());
@@ -1097,7 +1097,7 @@ DEF_OP(LRem) {
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
switch (OpSize) {
case 2: {
case IR::OpSize::i16Bit: {
uxth(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 16, 16);
sxth(EmitSize, TMP2, Divisor);
@@ -1105,7 +1105,7 @@ DEF_OP(LRem) {
msub(EmitSize, Dst, TMP3, TMP2, TMP1);
break;
}
case 4: {
case IR::OpSize::i32Bit: {
// TODO: 32-bit operation should be guaranteed not to leave garbage in the upper bits.
mov(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 32, 32);
@@ -1114,9 +1114,9 @@ DEF_OP(LRem) {
msub(EmitSize, Dst, TMP2, TMP3, TMP1);
break;
}
case 8: {
ARMEmitter::SingleUseForwardLabel Only64Bit {};
ARMEmitter::SingleUseForwardLabel LongDIVRet {};
case IR::OpSize::i64Bit: {
ARMEmitter::ForwardLabel Only64Bit {};
ARMEmitter::ForwardLabel LongDIVRet {};
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
@@ -1160,8 +1160,8 @@ DEF_OP(LRem) {
DEF_OP(LURem) {
auto Op = IROp->C<IR::IROp_LURem>();
const uint8_t OpSize = IROp->Size;
const auto EmitSize = OpSize >= 4 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto OpSize = IROp->Size;
const auto EmitSize = OpSize >= IR::OpSize::i32Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Dst = GetReg(Node);
const auto Upper = GetReg(Op->Upper.ID());
@@ -1171,14 +1171,14 @@ DEF_OP(LURem) {
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
switch (OpSize) {
case 2: {
case IR::OpSize::i16Bit: {
uxth(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 16, 16);
udiv(EmitSize, TMP2, TMP1, Divisor);
msub(EmitSize, Dst, TMP2, Divisor, TMP1);
break;
}
case 4: {
case IR::OpSize::i32Bit: {
// TODO: 32-bit operation should be guaranteed not to leave garbage in the upper bits.
mov(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 32, 32);
@@ -1186,9 +1186,9 @@ DEF_OP(LURem) {
msub(EmitSize, Dst, TMP2, Divisor, TMP1);
break;
}
case 8: {
ARMEmitter::SingleUseForwardLabel Only64Bit {};
ARMEmitter::SingleUseForwardLabel LongDIVRet {};
case IR::OpSize::i64Bit: {
ARMEmitter::ForwardLabel Only64Bit {};
ARMEmitter::ForwardLabel LongDIVRet {};
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
@@ -1238,30 +1238,30 @@ DEF_OP(Not) {
DEF_OP(Popcount) {
auto Op = IROp->C<IR::IROp_Popcount>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src.ID());
switch (OpSize) {
case 0x1:
case IR::OpSize::i8Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
// only use lowest byte
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
case 0x2:
case IR::OpSize::i16Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// only count two lowest bytes
addp(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D(), VTMP1.D());
break;
case 0x4:
case IR::OpSize::i32Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// fmov has zero extended, unused bytes are zero
addv(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
case 0x8:
case IR::OpSize::i64Bit:
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// fmov has zero extended, unused bytes are zero
@@ -1280,34 +1280,27 @@ DEF_OP(FindLSB) {
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src.ID());
if (IROp->Size != 8) {
ubfx(EmitSize, TMP1, Src, 0, IROp->Size * 8);
cmp(EmitSize, TMP1, 0);
rbit(EmitSize, TMP1, TMP1);
} else {
rbit(EmitSize, TMP1, Src);
cmp(EmitSize, Src, 0);
}
// We assume the source is nonzero, so we can just rbit+clz without worrying
// about upper garbage for smaller types.
rbit(EmitSize, TMP1, Src);
clz(EmitSize, Dst, TMP1);
csinv(EmitSize, Dst, Dst, ARMEmitter::Reg::zr, ARMEmitter::Condition::CC_NE);
}
DEF_OP(FindMSB) {
auto Op = IROp->C<IR::IROp_FindMSB>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i16Bit || OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit,
"Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src.ID());
movz(ARMEmitter::Size::i64Bit, TMP1, OpSize * 8 - 1);
movz(ARMEmitter::Size::i64Bit, TMP1, IR::OpSizeAsBits(OpSize) - 1);
if (OpSize == 2) {
if (OpSize == IR::OpSize::i16Bit) {
lsl(EmitSize, Dst, Src, 16);
orr(EmitSize, Dst, Dst, 0x8000);
clz(EmitSize, Dst, Dst);
} else {
clz(EmitSize, Dst, Src);
@@ -1318,9 +1311,10 @@ DEF_OP(FindMSB) {
DEF_OP(FindTrailingZeroes) {
auto Op = IROp->C<IR::IROp_FindTrailingZeroes>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i16Bit || OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit,
"Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
@@ -1328,7 +1322,7 @@ DEF_OP(FindTrailingZeroes) {
rbit(EmitSize, Dst, Src);
if (OpSize == 2) {
if (OpSize == IR::OpSize::i16Bit) {
// This orr does two things. First, if the (masked) source is zero, it
// reverses to zero in the top so it forces clz to return 16. Second, it
// ensures garbage in the upper bits of the source don't affect clz, because
@@ -1342,15 +1336,16 @@ DEF_OP(FindTrailingZeroes) {
DEF_OP(CountLeadingZeroes) {
auto Op = IROp->C<IR::IROp_CountLeadingZeroes>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i16Bit || OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit,
"Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src.ID());
if (OpSize == 2) {
if (OpSize == IR::OpSize::i16Bit) {
// Expressing as lsl+orr+clz clears away any garbage in the upper bits
// (alternatively could do uxth+clz+sub.. equal cost in total).
lsl(EmitSize, Dst, Src, 16);
@@ -1363,16 +1358,17 @@ DEF_OP(CountLeadingZeroes) {
DEF_OP(Rev) {
auto Op = IROp->C<IR::IROp_Rev>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i16Bit || OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit,
"Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src.ID());
rev(EmitSize, Dst, Src);
if (OpSize == 2) {
if (OpSize == IR::OpSize::i16Bit) {
lsr(EmitSize, Dst, Dst, 16);
}
}
@@ -1398,10 +1394,10 @@ DEF_OP(Bfi) {
mov(EmitSize, TMP1, SrcDst);
bfi(EmitSize, TMP1, Src, Op->lsb, Op->Width);
if (IROp->Size >= 4) {
if (IROp->Size >= IR::OpSize::i32Bit) {
mov(EmitSize, Dst, TMP1.R());
} else {
ubfx(EmitSize, Dst, TMP1, 0, IROp->Size * 8);
ubfx(EmitSize, Dst, TMP1, 0, IR::OpSizeAsBits(IROp->Size));
}
}
}
@@ -1432,8 +1428,8 @@ DEF_OP(Bfxil) {
DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
LOGMAN_THROW_AA_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
LOGMAN_THROW_AA_FMT(Op->Width != 0, "Invalid BFE width of 0");
LOGMAN_THROW_A_FMT(IROp->Size <= IR::OpSize::i64Bit, "OpSize is too large for BFE: {}", IROp->Size);
LOGMAN_THROW_A_FMT(Op->Width != 0, "Invalid BFE width of 0");
const auto EmitSize = ConvertSize(IROp);
const auto Dst = GetReg(Node);
@@ -1442,7 +1438,7 @@ DEF_OP(Bfe) {
if (Op->lsb == 0 && Op->Width == 32) {
mov(ARMEmitter::Size::i32Bit, Dst, Src);
} else if (Op->lsb == 0 && Op->Width == 64) {
LOGMAN_THROW_AA_FMT(IROp->Size == 8, "Must be 64-bit wide register");
LOGMAN_THROW_A_FMT(IROp->Size == IR::OpSize::i64Bit, "Must be 64-bit wide register");
mov(ARMEmitter::Size::i64Bit, Dst, Src);
} else {
ubfx(EmitSize, Dst, Src, Op->lsb, Op->Width);
@@ -1459,9 +1455,9 @@ DEF_OP(Sbfe) {
DEF_OP(Select) {
auto Op = IROp->C<IR::IROp_Select>();
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto EmitSize = ConvertSize(IROp);
const auto CompareEmitSize = Op->CompareSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto CompareEmitSize = Op->CompareSize == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
uint64_t Const;
auto cc = MapCC(Op->Cond);
@@ -1478,7 +1474,7 @@ DEF_OP(Select) {
} else if (IsFPR(Op->Cmp1.ID())) {
const auto Src1 = GetVReg(Op->Cmp1.ID());
const auto Src2 = GetVReg(Op->Cmp2.ID());
fcmp(Op->CompareSize == 8 ? ARMEmitter::ScalarRegSize::i64Bit : ARMEmitter::ScalarRegSize::i32Bit, Src1, Src2);
fcmp(Op->CompareSize == IR::OpSize::i64Bit ? ARMEmitter::ScalarRegSize::i64Bit : ARMEmitter::ScalarRegSize::i32Bit, Src1, Src2);
} else {
LOGMAN_MSG_A_FMT("Select: Expected GPR or FPR");
}
@@ -1487,7 +1483,7 @@ DEF_OP(Select) {
bool is_const_true = IsInlineConstant(Op->TrueVal, &const_true);
bool is_const_false = IsInlineConstant(Op->FalseVal, &const_false);
uint64_t all_ones = OpSize == 8 ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
uint64_t all_ones = OpSize == IR::OpSize::i64Bit ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
ARMEmitter::Register Dst = GetReg(Node);
@@ -1516,7 +1512,7 @@ DEF_OP(NZCVSelect) {
bool is_const_true = IsInlineConstant(Op->TrueVal, &const_true);
bool is_const_false = IsInlineConstant(Op->FalseVal, &const_false);
uint64_t all_ones = IROp->Size == 8 ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
uint64_t all_ones = IROp->Size == IR::OpSize::i64Bit ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
ARMEmitter::Register Dst = GetReg(Node);
@@ -1535,6 +1531,14 @@ DEF_OP(NZCVSelect) {
}
}
DEF_OP(NZCVSelectV) {
auto Op = IROp->C<IR::IROp_NZCVSelectV>();
auto cc = MapCC(Op->Cond);
const auto SubRegSize = ConvertSubRegSizePair248(IROp);
fcsel(SubRegSize.Scalar, GetVReg(Node), GetVReg(Op->TrueVal.ID()), GetVReg(Op->FalseVal.ID()), cc);
}
DEF_OP(NZCVSelectIncrement) {
auto Op = IROp->C<IR::IROp_NZCVSelectIncrement>();
@@ -1545,12 +1549,12 @@ DEF_OP(VExtractToGPR) {
const auto Op = IROp->C<IR::IROp_VExtractToGPR>();
const auto OpSize = IROp->Size;
constexpr auto AVXRegBitSize = Core::CPUState::XMM_AVX_REG_SIZE * 8;
[[maybe_unused]] constexpr auto AVXRegBitSize = Core::CPUState::XMM_AVX_REG_SIZE * 8;
constexpr auto SSERegBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
const auto ElementSizeBits = Op->Header.ElementSize * 8;
const auto ElementSizeBits = IR::OpSizeAsBits(Op->Header.ElementSize);
const auto Offset = ElementSizeBits * Op->Index;
const auto Is256Bit = Offset >= SSERegBitSize;
[[maybe_unused]] const auto Is256Bit = Offset >= SSERegBitSize;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetReg(Node);
@@ -1558,10 +1562,10 @@ DEF_OP(VExtractToGPR) {
const auto PerformMove = [&](const ARMEmitter::VRegister reg, int index) {
switch (OpSize) {
case 1: umov<ARMEmitter::SubRegSize::i8Bit>(Dst, Vector, index); break;
case 2: umov<ARMEmitter::SubRegSize::i16Bit>(Dst, Vector, index); break;
case 4: umov<ARMEmitter::SubRegSize::i32Bit>(Dst, Vector, index); break;
case 8: umov<ARMEmitter::SubRegSize::i64Bit>(Dst, Vector, index); break;
case IR::OpSize::i8Bit: umov<ARMEmitter::SubRegSize::i8Bit>(Dst, Vector, index); break;
case IR::OpSize::i16Bit: umov<ARMEmitter::SubRegSize::i16Bit>(Dst, Vector, index); break;
case IR::OpSize::i32Bit: umov<ARMEmitter::SubRegSize::i32Bit>(Dst, Vector, index); break;
case IR::OpSize::i64Bit: umov<ARMEmitter::SubRegSize::i64Bit>(Dst, Vector, index); break;
default: LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", OpSize); break;
}
};
@@ -1572,8 +1576,8 @@ DEF_OP(VExtractToGPR) {
// when acting on larger register sizes.
PerformMove(Vector, Op->Index);
} else {
LOGMAN_THROW_AA_FMT(Is256Bit, "Can't perform 256-bit extraction with op side: {}", OpSize);
LOGMAN_THROW_AA_FMT(Offset < AVXRegBitSize, "Trying to extract element outside bounds of register. Offset={}, Index={}", Offset, Op->Index);
LOGMAN_THROW_A_FMT(Is256Bit, "Can't perform 256-bit extraction with op side: {}", OpSize);
LOGMAN_THROW_A_FMT(Offset < AVXRegBitSize, "Trying to extract element outside bounds of register. Offset={}, Index={}", Offset, Op->Index);
// We need to use the upper 128-bit lane, so lets move it down.
// Inverting our dedicated predicate for 128-bit operations selects
@@ -1586,10 +1590,10 @@ DEF_OP(VExtractToGPR) {
// upper half of the vector.
const auto SanitizedIndex = [OpSize, Op] {
switch (OpSize) {
case 1: return Op->Index - 16;
case 2: return Op->Index - 8;
case 4: return Op->Index - 4;
case 8: return Op->Index - 2;
case IR::OpSize::i8Bit: return Op->Index - 16;
case IR::OpSize::i16Bit: return Op->Index - 8;
case IR::OpSize::i32Bit: return Op->Index - 4;
case IR::OpSize::i64Bit: return Op->Index - 2;
default: LOGMAN_MSG_A_FMT("Unhandled OpSize: {}", OpSize); return 0;
}
}();
@@ -1605,7 +1609,7 @@ DEF_OP(Float_ToGPR_ZS) {
ARMEmitter::Register Dst = GetReg(Node);
ARMEmitter::VRegister Src = GetVReg(Op->Scalar.ID());
if (Op->SrcElementSize == 8) {
if (Op->SrcElementSize == IR::OpSize::i64Bit) {
fcvtzs(ConvertSize(IROp), Dst, Src.D());
} else {
fcvtzs(ConvertSize(IROp), Dst, Src.S());
@@ -1618,7 +1622,7 @@ DEF_OP(Float_ToGPR_S) {
ARMEmitter::Register Dst = GetReg(Node);
ARMEmitter::VRegister Src = GetVReg(Op->Scalar.ID());
if (Op->SrcElementSize == 8) {
if (Op->SrcElementSize == IR::OpSize::i64Bit) {
frinti(VTMP1.D(), Src.D());
fcvtzs(ConvertSize(IROp), Dst, VTMP1.D());
} else {
@@ -1629,7 +1633,7 @@ DEF_OP(Float_ToGPR_S) {
DEF_OP(FCmp) {
auto Op = IROp->C<IR::IROp_FCmp>();
const auto EmitSubSize = Op->ElementSize == 8 ? ARMEmitter::ScalarRegSize::i64Bit : ARMEmitter::ScalarRegSize::i32Bit;
const auto EmitSubSize = Op->ElementSize == IR::OpSize::i64Bit ? ARMEmitter::ScalarRegSize::i64Bit : ARMEmitter::ScalarRegSize::i32Bit;
ARMEmitter::VRegister Scalar1 = GetVReg(Op->Scalar1.ID());
ARMEmitter::VRegister Scalar2 = GetVReg(Op->Scalar2.ID());
@@ -6,7 +6,7 @@ desc: relocation logic of the arm64 splatter backend
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
#include <FEXCore/Core/Thunks.h>
@@ -86,7 +86,7 @@ bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uin
size_t DataIndex {};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation* Reloc = reinterpret_cast<const FEXCore::CPU::Relocation*>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_AA_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
@@ -7,13 +7,13 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
LOGMAN_THROW_AA_FMT(IROp->ElementSize == 4 || IROp->ElementSize == 8, "Wrong element size");
LOGMAN_THROW_A_FMT(IROp->ElementSize == IR::OpSize::i32Bit || IROp->ElementSize == IR::OpSize::i64Bit, "Wrong element size");
// Size is the size of each pair element
auto Dst0 = GetReg(Op->OutLo.ID());
auto Dst1 = GetReg(Op->OutHi.ID());
@@ -23,7 +23,7 @@ DEF_OP(CASPair) {
auto Desired1 = GetReg(Op->DesiredHi.ID());
auto MemSrc = GetReg(Op->Addr.ID());
const auto EmitSize = IROp->ElementSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto EmitSize = IROp->ElementSize == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
if (CTX->HostFeatures.SupportsAtomics) {
// RA has heuristics to try to pair sources, but we need to handle the cases
// where they fail. We do so by moving to temporaries. Note we use 64-bit
@@ -61,8 +61,8 @@ DEF_OP(CASPair) {
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::SingleUseForwardLabel LoopNotExpected;
ARMEmitter::SingleUseForwardLabel LoopExpected;
ARMEmitter::ForwardLabel LoopNotExpected;
ARMEmitter::ForwardLabel LoopExpected;
Bind(&LoopTop);
// This instruction sequence must be synced with HandleCASPAL_Armv8.
@@ -108,13 +108,13 @@ DEF_OP(CAS) {
mov(EmitSize, GetReg(Node), TMP2.R());
} else {
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::SingleUseForwardLabel LoopNotExpected;
ARMEmitter::SingleUseForwardLabel LoopExpected;
ARMEmitter::ForwardLabel LoopNotExpected;
ARMEmitter::ForwardLabel LoopExpected;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
if (IROp->Size == 1) {
if (IROp->Size == IR::OpSize::i8Bit) {
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTB, 0);
} else if (IROp->Size == 2) {
} else if (IROp->Size == IR::OpSize::i16Bit) {
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTH, 0);
} else {
cmp(EmitSize, TMP2, Expected);
@@ -273,18 +273,21 @@ DEF_OP(AtomicNeg) {
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(
OpSize == IR::OpSize::i64Bit || OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i16Bit || OpSize == IR::OpSize::i8Bit, "Unexpecte"
"d CAS "
"size");
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
const auto SubEmitSize = OpSize == IR::OpSize::i64Bit ? ARMEmitter::SubRegSize::i64Bit :
OpSize == IR::OpSize::i32Bit ? ARMEmitter::SubRegSize::i32Bit :
OpSize == IR::OpSize::i16Bit ? ARMEmitter::SubRegSize::i16Bit :
OpSize == IR::OpSize::i8Bit ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldswpal(SubEmitSize, Src, GetReg(Node), MemSrc);
@@ -294,7 +297,7 @@ DEF_OP(AtomicSwap) {
ldaxr(SubEmitSize, TMP2, MemSrc);
stlxr(SubEmitSize, TMP4, Src, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
ubfm(EmitSize, GetReg(Node), TMP2, 0, OpSize * 8 - 1);
ubfm(EmitSize, GetReg(Node), TMP2, 0, IR::OpSizeAsBits(OpSize) - 1);
}
}
@@ -9,7 +9,7 @@ $end_info$
#include "FEXCore/IR/IR.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
#include <FEXCore/Core/Thunks.h>
#include <FEXCore/Core/X86Enums.h>
@@ -53,14 +53,14 @@ DEF_OP(ExitFunction) {
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef _M_ARM_64EC
if (RtlIsEcCode(NewRIP)) {
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
LoadConstant(ARMEmitter::Size::i64Bit, EC_CALL_CHECKER_PC_REG, NewRIP);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
} else {
#endif
ARMEmitter::SingleUseForwardLabel l_BranchHost;
ARMEmitter::ForwardLabel l_BranchHost;
ldr(TMP1, &l_BranchHost);
blr(TMP1);
@@ -72,7 +72,7 @@ DEF_OP(ExitFunction) {
#endif
} else {
ARMEmitter::SingleUseForwardLabel FullLookup;
ARMEmitter::ForwardLabel FullLookup;
auto RipReg = GetReg(Op->NewRIP.ID());
// L1 Cache
@@ -117,7 +117,7 @@ DEF_OP(CondJump) {
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
auto Reg = GetReg(Op->Cmp1.ID());
const auto Size = Op->CompareSize == 4 ? ARMEmitter::Size::i32Bit : ARMEmitter::Size::i64Bit;
const auto Size = Op->CompareSize == IR::OpSize::i32Bit ? ARMEmitter::Size::i32Bit : ARMEmitter::Size::i64Bit;
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
LOGMAN_THROW_A_FMT(isConst, "CondJump: Expected constant source");
@@ -5,7 +5,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
@@ -15,18 +15,18 @@ DEF_OP(VInsGPR) {
const auto DestIdx = Op->DestIdx;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto SubEmitSize = ConvertSubRegSize8(IROp);
const auto ElementsPer128Bit = 16 / ElementSize;
const auto ElementsPer128Bit = IR::NumElements(IR::OpSize::i128Bit, ElementSize);
const auto Dst = GetVReg(Node);
const auto DestVector = GetVReg(Op->DestVector.ID());
const auto Src = GetReg(Op->Src.ID());
if (HostSupportsSVE256 && Is256Bit) {
const auto ElementSizeBits = ElementSize * 8;
const auto ElementSizeBits = IR::OpSizeAsBits(ElementSize);
const auto Offset = ElementSizeBits * DestIdx;
const auto SSEBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
@@ -90,16 +90,16 @@ DEF_OP(VCastFromGPR) {
auto Src = GetReg(Op->Src.ID());
switch (Op->Header.ElementSize) {
case 1:
case IR::OpSize::i8Bit:
uxtb(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 2:
case IR::OpSize::i16Bit:
uxth(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 4: fmov(ARMEmitter::Size::i32Bit, Dst.S(), Src); break;
case 8: fmov(ARMEmitter::Size::i64Bit, Dst.D(), Src); break;
case IR::OpSize::i32Bit: fmov(ARMEmitter::Size::i32Bit, Dst.S(), Src); break;
case IR::OpSize::i64Bit: fmov(ARMEmitter::Size::i64Bit, Dst.D(), Src); break;
default: LOGMAN_MSG_A_FMT("Unknown castGPR element size: {}", Op->Header.ElementSize);
}
}
@@ -111,7 +111,7 @@ DEF_OP(VDupFromGPR) {
const auto Dst = GetVReg(Node);
const auto Src = GetReg(Op->Src.ID());
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto SubEmitSize = ConvertSubRegSize8(IROp);
@@ -126,8 +126,8 @@ DEF_OP(VDupFromGPR) {
DEF_OP(Float_FromGPR_S) {
const auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
const uint16_t ElementSize = Op->Header.ElementSize;
const uint16_t Conv = (ElementSize << 8) | Op->SrcElementSize;
const uint16_t ElementSize = IR::OpSizeToSize(Op->Header.ElementSize);
const uint16_t Conv = (ElementSize << 8) | IR::OpSizeToSize(Op->SrcElementSize);
auto Dst = GetVReg(Node);
auto Src = GetReg(Op->Src.ID());
@@ -165,7 +165,7 @@ DEF_OP(Float_FromGPR_S) {
DEF_OP(Float_FToF) {
auto Op = IROp->C<IR::IROp_Float_FToF>();
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (IR::OpSizeToSize(Op->Header.ElementSize) << 8) | IR::OpSizeToSize(Op->SrcElementSize);
auto Dst = GetVReg(Node);
auto Src = GetVReg(Op->Scalar.ID());
@@ -205,7 +205,7 @@ DEF_OP(Vector_SToF) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
@@ -215,15 +215,15 @@ DEF_OP(Vector_SToF) {
scvtf(Dst.Z(), SubEmitSize, Mask.Merging(), Vector.Z(), SubEmitSize);
} else {
if (OpSize == ElementSize) {
if (ElementSize == 8) {
if (ElementSize == IR::OpSize::i64Bit) {
scvtf(ARMEmitter::ScalarRegSize::i64Bit, Dst.D(), Vector.D());
} else if (ElementSize == 4) {
} else if (ElementSize == IR::OpSize::i32Bit) {
scvtf(ARMEmitter::ScalarRegSize::i32Bit, Dst.S(), Vector.S());
} else {
scvtf(ARMEmitter::ScalarRegSize::i16Bit, Dst.H(), Vector.H());
}
} else {
if (OpSize == 8) {
if (OpSize == IR::OpSize::i64Bit) {
scvtf(SubEmitSize, Dst.D(), Vector.D());
} else {
scvtf(SubEmitSize, Dst.Q(), Vector.Q());
@@ -238,7 +238,7 @@ DEF_OP(Vector_FToZS) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
@@ -248,15 +248,15 @@ DEF_OP(Vector_FToZS) {
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Vector.Z(), SubEmitSize);
} else {
if (OpSize == ElementSize) {
if (ElementSize == 8) {
if (ElementSize == IR::OpSize::i64Bit) {
fcvtzs(ARMEmitter::ScalarRegSize::i64Bit, Dst.D(), Vector.D());
} else if (ElementSize == 4) {
} else if (ElementSize == IR::OpSize::i32Bit) {
fcvtzs(ARMEmitter::ScalarRegSize::i32Bit, Dst.S(), Vector.S());
} else {
fcvtzs(ARMEmitter::ScalarRegSize::i16Bit, Dst.H(), Vector.H());
}
} else {
if (OpSize == 8) {
if (OpSize == IR::OpSize::i64Bit) {
fcvtzs(SubEmitSize, Dst.D(), Vector.D());
} else {
fcvtzs(SubEmitSize, Dst.Q(), Vector.Q());
@@ -269,7 +269,7 @@ DEF_OP(Vector_FToS) {
const auto Op = IROp->C<IR::IROp_Vector_FToS>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto SubEmitSize = ConvertSubRegSize248(IROp);
@@ -284,7 +284,7 @@ DEF_OP(Vector_FToS) {
} else {
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (OpSize == 8) {
if (OpSize == IR::OpSize::i64Bit) {
frinti(SubEmitSize, Dst.D(), Vector.D());
fcvtzs(SubEmitSize, Dst.D(), Dst.D());
} else {
@@ -300,10 +300,10 @@ DEF_OP(Vector_FToF) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Conv = (ElementSize << 8) | Op->SrcElementSize;
const auto Conv = (IR::OpSizeToSize(ElementSize) << 8) | IR::OpSizeToSize(Op->SrcElementSize);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -403,7 +403,7 @@ DEF_OP(Vector_FToI) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
@@ -427,15 +427,15 @@ DEF_OP(Vector_FToI) {
// frinti having AdvSIMD, AdvSIMD scalar, and an SVE version),
// we can't just use a lambda without some seriously ugly casting.
// This is fairly self-contained otherwise.
#define ROUNDING_FN(name) \
if (ElementSize == 2) { \
name(Dst.H(), Vector.H()); \
} else if (ElementSize == 4) { \
name(Dst.S(), Vector.S()); \
} else if (ElementSize == 8) { \
name(Dst.D(), Vector.D()); \
} else { \
FEX_UNREACHABLE; \
#define ROUNDING_FN(name) \
if (ElementSize == IR::OpSize::i16Bit) { \
name(Dst.H(), Vector.H()); \
} else if (ElementSize == IR::OpSize::i32Bit) { \
name(Dst.S(), Vector.S()); \
} else if (ElementSize == IR::OpSize::i64Bit) { \
name(Dst.D(), Vector.D()); \
} else { \
FEX_UNREACHABLE; \
}
switch (Op->Round) {
@@ -464,7 +464,7 @@ DEF_OP(Vector_F64ToI32) {
const auto OpSize = IROp->Size;
const auto Round = Op->Round;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || (Is256Bit && HostSupportsSVE256), "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
@@ -5,7 +5,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
@@ -17,14 +17,14 @@ DEF_OP(VAESImc) {
DEF_OP(VAESEnc) {
const auto Op = IROp->C<IR::IROp_VAESEnc>();
const auto OpSize = IROp->Size;
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
@@ -42,14 +42,14 @@ DEF_OP(VAESEnc) {
DEF_OP(VAESEncLast) {
const auto Op = IROp->C<IR::IROp_VAESEncLast>();
const auto OpSize = IROp->Size;
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
@@ -65,14 +65,14 @@ DEF_OP(VAESEncLast) {
DEF_OP(VAESDec) {
const auto Op = IROp->C<IR::IROp_VAESDec>();
const auto OpSize = IROp->Size;
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
@@ -90,14 +90,14 @@ DEF_OP(VAESDec) {
DEF_OP(VAESDecLast) {
const auto Op = IROp->C<IR::IROp_VAESDecLast>();
const auto OpSize = IROp->Size;
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
@@ -152,10 +152,10 @@ DEF_OP(CRC32) {
const auto Src2 = GetReg(Op->Src2.ID());
switch (Op->SrcSize) {
case 1: crc32cb(Dst.W(), Src1.W(), Src2.W()); break;
case 2: crc32ch(Dst.W(), Src1.W(), Src2.W()); break;
case 4: crc32cw(Dst.W(), Src1.W(), Src2.W()); break;
case 8: crc32cx(Dst.X(), Src1.X(), Src2.X()); break;
case IR::OpSize::i8Bit: crc32cb(Dst.W(), Src1.W(), Src2.W()); break;
case IR::OpSize::i16Bit: crc32ch(Dst.W(), Src1.W(), Src2.W()); break;
case IR::OpSize::i32Bit: crc32cw(Dst.W(), Src1.W(), Src2.W()); break;
case IR::OpSize::i64Bit: crc32cx(Dst.X(), Src1.X(), Src2.X()); break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
}
}
@@ -169,6 +169,85 @@ DEF_OP(VSha1H) {
sha1h(Dst.S(), Src.S());
}
DEF_OP(VSha1C) {
auto Op = IROp->C<IR::IROp_VSha1C>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
if (Dst == Src1) {
sha1c(Dst, Src2.S(), Src3);
} else if (Dst != Src2 && Dst != Src3) {
mov(Dst.Q(), Src1.Q());
sha1c(Dst, Src2.S(), Src3);
} else {
mov(VTMP1.Q(), Src1.Q());
sha1c(VTMP1, Src2.S(), Src3);
mov(Dst.Q(), VTMP1.Q());
}
}
DEF_OP(VSha1M) {
auto Op = IROp->C<IR::IROp_VSha1M>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
if (Dst == Src1) {
sha1m(Dst, Src2.S(), Src3);
} else if (Dst != Src2 && Dst != Src3) {
mov(Dst.Q(), Src1.Q());
sha1m(Dst, Src2.S(), Src3);
} else {
mov(VTMP1.Q(), Src1.Q());
sha1m(VTMP1, Src2.S(), Src3);
mov(Dst.Q(), VTMP1.Q());
}
}
DEF_OP(VSha1P) {
auto Op = IROp->C<IR::IROp_VSha1P>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
if (Dst == Src1) {
sha1p(Dst, Src2.S(), Src3);
} else if (Dst != Src2 && Dst != Src3) {
mov(Dst.Q(), Src1.Q());
sha1p(Dst, Src2.S(), Src3);
} else {
mov(VTMP1.Q(), Src1.Q());
sha1p(VTMP1, Src2.S(), Src3);
mov(Dst.Q(), VTMP1.Q());
}
}
DEF_OP(VSha1SU1) {
auto Op = IROp->C<IR::IROp_VSha1SU1>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
if (Dst == Src1) {
sha1su1(Dst, Src2);
} else if (Dst != Src2) {
mov(Dst.Q(), Src1.Q());
sha1su1(Dst, Src2);
} else {
mov(VTMP1.Q(), Src1.Q());
sha1su1(VTMP1, Src2);
mov(Dst.Q(), VTMP1.Q());
}
}
DEF_OP(VSha256U0) {
auto Op = IROp->C<IR::IROp_VSha256U0>();
@@ -185,15 +264,32 @@ DEF_OP(VSha256U0) {
}
}
DEF_OP(PCLMUL) {
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
DEF_OP(VSha256U1) {
auto Op = IROp->C<IR::IROp_VSha256U1>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
if (Dst != Src1 && Dst != Src1) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
sha256su1(Dst, Src1, Src2);
} else {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
sha256su1(VTMP1, Src1, Src2);
mov(Dst.Q(), VTMP1.Q());
}
}
DEF_OP(PCLMUL) {
const auto Op = IROp->C<IR::IROp_PCLMUL>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
switch (Op->Selector) {
case 0b00000000: pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D()); break;
@@ -11,16 +11,18 @@ desc: Main glue logic of the arm64 splatter backend
$end_info$
*/
#include "Common/SoftFloat.h"
#include "FEXCore/Utils/Telemetry.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Utils/MemberFunctionToPointer.h"
#include "Utils/variable_length_integer.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
@@ -35,6 +37,7 @@ $end_info$
#include <stdio.h>
#include <unistd.h>
#include <string.h>
#include <limits>
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
@@ -85,16 +88,13 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
} else {
auto FillF80Result = [&]() {
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
mov(TMP2, ARMEmitter::XReg::x1);
mov(VTMP1.Q(), ARMEmitter::VReg::v0.Q());
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
eor(Dst.Q(), Dst.Q(), Dst.Q());
ins(ARMEmitter::SubRegSize::i64Bit, Dst, 0, TMP1);
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, TMP2);
mov(Dst.Q(), VTMP1.Q());
};
auto FillF64Result = [&]() {
@@ -118,52 +118,16 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
};
switch (Info.ABI) {
case FABI_F80_I16_F32: {
case FABI_F80_I16_F32_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
fmov(ARMEmitter::SReg::s0, Src1.S());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, float>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
FillF80Result();
} break;
case FABI_F80_I16_F64: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
mov(ARMEmitter::DReg::d0, Src1.D());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, double>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
FillF80Result();
} break;
case FABI_F80_I16_I16:
case FABI_F80_I16_I32: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetReg(IROp->Args[0].ID());
if (Info.ABI == FABI_F80_I16_I16) {
sxth(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
} else {
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint32_t>(ARMEmitter::Reg::r2);
GenerateIndirectRuntimeCall<FEXCore::VectorRegType, uint16_t, float, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
@@ -171,20 +135,59 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
FillF80Result();
} break;
case FABI_F32_I16_F80: {
case FABI_F80_I16_F64_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
mov(ARMEmitter::DReg::d0, Src1.D());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<FEXCore::VectorRegType, uint16_t, double, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
FillF80Result();
} break;
case FABI_F80_I16_I16_PTR:
case FABI_F80_I16_I32_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetReg(IROp->Args[0].ID());
if (Info.ABI == FABI_F80_I16_I16_PTR) {
sxth(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
} else {
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x2, STATE);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<FEXCore::VectorRegType, uint16_t, uint32_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
FillF80Result();
} break;
case FABI_F32_I16_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<float, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<float, uint16_t, FEXCore::VectorRegType, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r3);
blr(ARMEmitter::Reg::r2);
}
if (!TMP_ABIARGS) {
@@ -196,43 +199,45 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
fmov(Dst.S(), VTMP1.S());
} break;
case FABI_F64_I16_F80: {
case FABI_F64_I16_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<double, uint16_t, FEXCore::VectorRegType, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r3);
blr(ARMEmitter::Reg::r2);
}
FillF64Result();
} break;
case FABI_F64_I16_F64: {
case FABI_F64_I16_F64_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
mov(ARMEmitter::DReg::d0, Src1.D());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double>(ARMEmitter::Reg::r1);
GenerateIndirectRuntimeCall<double, uint16_t, double, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r1);
blr(ARMEmitter::Reg::r2);
}
FillF64Result();
} break;
case FABI_F64_I16_F64_F64: {
case FABI_F64_I16_F64_F64_PTR: {
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
@@ -247,30 +252,31 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double, double>(ARMEmitter::Reg::r1);
GenerateIndirectRuntimeCall<double, uint16_t, double, double, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r1);
blr(ARMEmitter::Reg::r2);
}
FillF64Result();
} break;
case FABI_I16_I16_F80: {
case FABI_I16_I16_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<uint32_t, uint16_t, FEXCore::VectorRegType, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r3);
blr(ARMEmitter::Reg::r2);
}
if (!TMP_ABIARGS) {
@@ -281,38 +287,38 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
const auto Dst = GetReg(Node);
sxth(ARMEmitter::Size::i64Bit, Dst, TMP1);
} break;
case FABI_I32_I16_F80: {
case FABI_I32_I16_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<uint32_t, uint16_t, FEXCore::VectorRegType, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r3);
blr(ARMEmitter::Reg::r2);
}
FillI32Result();
} break;
case FABI_I64_I16_F80: {
case FABI_I64_I16_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<uint64_t, uint16_t, FEXCore::VectorRegType, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r3);
blr(ARMEmitter::Reg::r2);
}
if (!TMP_ABIARGS) {
@@ -323,24 +329,28 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
} break;
case FABI_I64_I16_F80_F80: {
case FABI_I64_I16_F80_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r4, Src2, 4);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint16_t, uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r5);
if (!TMP_ABIARGS) {
mov(VTMP1.Q(), Src1.Q());
mov(ARMEmitter::VReg::v1.Q(), Src2.Q());
mov(ARMEmitter::VReg::v0.Q(), VTMP1.Q());
} else {
blr(ARMEmitter::Reg::r5);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
mov(ARMEmitter::VReg::v1.Q(), Src2.Q());
}
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint16_t, FEXCore::VectorRegType, FEXCore::VectorRegType, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
if (!TMP_ABIARGS) {
@@ -351,42 +361,47 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
} break;
case FABI_F80_I16_F80: {
case FABI_F80_I16_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<FEXCore::VectorRegType, uint16_t, FEXCore::VectorRegType, uint64_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r3);
blr(ARMEmitter::Reg::r2);
}
FillF80Result();
} break;
case FABI_F80_I16_F80_F80: {
case FABI_F80_I16_F80_F80_PTR: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r4, Src2, 4);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r5);
if (!TMP_ABIARGS) {
mov(VTMP1.Q(), Src1.Q());
mov(ARMEmitter::VReg::v1.Q(), Src2.Q());
mov(ARMEmitter::VReg::v0.Q(), VTMP1.Q());
} else {
blr(ARMEmitter::Reg::r5);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
mov(ARMEmitter::VReg::v1.Q(), Src2.Q());
}
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<FEXCore::VectorRegType, uint16_t, FEXCore::VectorRegType, FEXCore::VectorRegType, uint64_t>(
ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
FillF80Result();
@@ -428,7 +443,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
FillI32Result();
} break;
case FABI_I32_I128_I128_I16: {
case FABI_I32_V128_V128_I16: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
@@ -437,19 +452,22 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
const auto Src2 = GetVReg(Op->RHS.ID());
const auto Control = Op->Control;
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r0, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 1);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src2, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 1);
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r4, Control);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t, uint64_t, uint64_t, uint16_t>(ARMEmitter::Reg::r5);
if (!TMP_ABIARGS) {
mov(VTMP1.Q(), Src1.Q());
mov(ARMEmitter::VReg::v1.Q(), Src2.Q());
mov(ARMEmitter::VReg::v0.Q(), VTMP1.Q());
} else {
blr(ARMEmitter::Reg::r5);
mov(ARMEmitter::VReg::v0.Q(), Src1.Q());
mov(ARMEmitter::VReg::v1.Q(), Src2.Q());
}
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r0, Control);
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, FEXCore::VectorRegType, FEXCore::VectorRegType, uint16_t>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
FillI32Result();
@@ -464,12 +482,11 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
uintptr_t branch = (uintptr_t)(Record)-8;
ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
ARMEmitter::SingleUseForwardLabel l_BranchHost;
ARMEmitter::ForwardLabel l_BranchHost;
emit.ldr(TMP1, &l_BranchHost);
emit.blr(TMP1);
emit.Bind(&l_BranchHost);
@@ -484,11 +501,16 @@ static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
bool TFSet = Thread->CurrentFrame->State.flags[X86State::RFLAG_TF_RAW_LOC];
uintptr_t HostCode {};
auto GuestRip = Record->GuestRIP;
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!TFSet) {
HostCode = Thread->LookupCache->FindBlock(GuestRip);
}
if (!HostCode) {
if (TFSet || !HostCode) {
// If TF is set, the cache must be skipped as different code needs to be generated.
Frame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
}
@@ -626,8 +648,8 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
uint64_t Mask = ~0ULL;
uint8_t OpSize = OpHeader->Size;
if (OpSize == 4) {
const auto Size = OpHeader->Size;
if (Size == IR::OpSize::i32Bit) {
Mask = 0xFFFF'FFFFULL;
}
*Value = (Entry + Op->Offset) & Mask;
@@ -654,8 +676,69 @@ bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData) {
void Arm64JITCore::EmitInterruptChecks(bool CheckTF) {
if (CheckTF) {
ARMEmitter::ForwardLabel l_TFUnset;
ARMEmitter::ForwardLabel l_TFBlocked;
// Note that this needs to be before the below suspend checks, as X86 checks this flag immediately after executing an instruction.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbz(ARMEmitter::Size::i32Bit, TMP1, &l_TFUnset);
// X86 semantically checks TF after executing each instruction, so e.g. setting a context with TF set will execute a single instruction
// and then raise an exception. However on the FEX side this is simpler to implement by checking at the start of each instruction, handle this by having bit 1 being unset in the flag state indicate that TF is blocked for a single instruction.
tbz(TMP1, 1, &l_TFBlocked);
// Block TF for a single instruction when the frontend jumps to a new context by unsetting bit 1.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
and_(ARMEmitter::Size::i32Bit, TMP1, TMP1, ~(1 << 1));
strb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
Core::CpuStateFrame::SynchronousFaultDataStruct State = {
.FaultToTopAndGeneratedException = 1,
.Signal = Core::FAULT_SIGTRAP,
.TrapNo = X86State::X86_TRAPNO_DB,
.si_code = 2,
.err_code = 0,
};
uint64_t Constant {};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Constant);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
Bind(&l_TFBlocked);
// If TF was blocked for this instruction, unblock it for the next.
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 0b11);
strb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
Bind(&l_TFUnset);
}
if (CTX->Config.NeedsPendingInterruptFaultCheck) {
// Trigger a fault if there are any pending interrupts
// Used only for suspend on WIN32 at the moment
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
}
#ifdef _M_ARM_64EC
static constexpr uint16_t SuspendMagic {0xCAFE};
ldr(TMP2.W(), STATE_PTR(CpuStateFrame, SuspendDoorbell));
ARMEmitter::ForwardLabel l_NoSuspend;
cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
brk(SuspendMagic);
Bind(&l_NoSuspend);
#endif
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, const FEXCore::IR::RegisterAllocationData* RAData,
bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
JumpTargets.clear();
@@ -711,22 +794,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore
adr(TMP1, &JITCodeHeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
if (CTX->Config.NeedsPendingInterruptFaultCheck) {
// Trigger a fault if there are any pending interrupts
// Used only for suspend on WIN32 at the moment
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
}
#ifdef _M_ARM_64EC
static constexpr uint16_t SuspendMagic {0xCAFE};
ldr(TMP2.W(), STATE_PTR(CpuStateFrame, SuspendDoorbell));
ARMEmitter::SingleUseForwardLabel l_NoSuspend;
cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
brk(SuspendMagic);
Bind(&l_NoSuspend);
#endif
EmitInterruptChecks(CheckTF);
SpillSlots = RAData->SpillSlots();
@@ -747,7 +815,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore
using namespace FEXCore::IR;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = GetCursorAddress<uint8_t*>();
@@ -800,34 +868,61 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
auto JITRIPEntriesLocation = GetCursorAddress<uint8_t*>();
auto JITRIPEntries = GetCursorAddress<JITRIPReconstructEntries*>();
// Entries that live after the JITCodeTail.
// These entries correlate JIT code regions with guest RIP regions.
// Using these entries FEX is able to reconstruct the guest RIP accurately when an instruction cause a signal fault.
// Packed using two variable length integer entries to ensure the size isn't too large.
// These smaller sizes means that each entry is relative to each other instead of absolute offset from the start of the JIT block.
// When reconstructing the RIP, each entry must be walked linearly and accumulated with the previous entries.
// This is a trade-off between compression inside the JIT code space and execution time when reconstruction the RIP.
// RIP reconstruction when faulting is less likely so we are requiring the accumulation.
//
// struct {
// // The Host PC offset from the previous entry.
// FEXCore::Utils::vl64 HostPCOffset;
// // How much to offset the RIP from the previous entry.
// FEXCore::Utils::vl64 GuestRIPOffset;
// };
CursorIncrement(sizeof(JITRIPReconstructEntries) * DebugData->GuestOpcodes.size());
auto JITRIPEntriesBegin = GetCursorAddress<uint8_t*>();
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
JITBlockTail->GuestSize = Size;
JITBlockTail->SingleInst = SingleInst;
JITBlockTail->SpinLockFutex = 0;
auto JITRIPEntriesLocation = JITRIPEntriesBegin;
{
// Store the RIP entries.
JITBlockTail->NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail->OffsetToRIPEntries = JITRIPEntriesLocation - JITBlockTailLocation;
JITBlockTail->OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
uintptr_t CurrentRIPOffset = 0;
uint64_t CurrentPCOffset = 0;
for (size_t i = 0; i < DebugData->GuestOpcodes.size(); i++) {
const auto& GuestOpcode = DebugData->GuestOpcodes[i];
auto& RIPEntry = JITRIPEntries[i];
RIPEntry.HostPCOffset = GuestOpcode.HostEntryOffset - CurrentPCOffset;
RIPEntry.GuestRIPOffset = GuestOpcode.GuestEntryOffset - CurrentRIPOffset;
int64_t HostPCOffset = GuestOpcode.HostEntryOffset - CurrentPCOffset;
int64_t GuestRIPOffset = GuestOpcode.GuestEntryOffset - CurrentRIPOffset;
size_t Size = FEXCore::Utils::vl64::Encode(JITRIPEntriesLocation, HostPCOffset);
JITRIPEntriesLocation += Size;
Size = FEXCore::Utils::vl64::Encode(JITRIPEntriesLocation, GuestRIPOffset);
JITRIPEntriesLocation += Size;
CurrentPCOffset = GuestOpcode.HostEntryOffset;
CurrentRIPOffset = GuestOpcode.GuestEntryOffset;
}
}
CursorIncrement(JITRIPEntriesLocation - JITRIPEntriesBegin);
Align();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
@@ -839,7 +934,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore
#ifdef VIXL_DISASSEMBLER
if (Disassemble() & FEXCore::Config::Disassemble::STATS) {
auto HeaderOp = IR->GetHeader();
LOGMAN_THROW_AA_FMT(HeaderOp->Header.Op == IR::OP_IRHEADER, "First op wasn't IRHeader");
LOGMAN_THROW_A_FMT(HeaderOp->Header.Op == IR::OP_IRHEADER, "First op wasn't IRHeader");
LogMan::Msg::IFmt("RIP: 0x{:x}", Entry);
LogMan::Msg::IFmt("Guest Code instructions: {}", HeaderOp->NumHostInstructions);
@@ -38,8 +38,9 @@ public:
~Arm64JITCore() override;
[[nodiscard]]
CPUBackend::CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData) override;
CPUBackend::CompiledCode
CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData, bool CheckTF) override;
void ClearCache() override;
@@ -68,7 +69,7 @@ private:
ARMEmitter::Register GetReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
LOGMAN_THROW_A_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::GPRFixedClass.Val) {
return StaticRegisters[Reg.Reg];
@@ -83,7 +84,7 @@ private:
ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
LOGMAN_THROW_A_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::FPRFixedClass.Val) {
return StaticFPRegisters[Reg.Reg];
@@ -110,7 +111,7 @@ private:
ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
uint64_t Const;
if (IsInlineConstant(Src, &Const)) {
LOGMAN_THROW_AA_FMT(Const == 0, "Only valid constant");
LOGMAN_THROW_A_FMT(Const == 0, "Only valid constant");
return ARMEmitter::Reg::zr;
} else {
return GetReg(Src.ID());
@@ -129,23 +130,25 @@ private:
[[nodiscard]]
ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
return Op->Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
return Op->Size == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
}
[[nodiscard]]
ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->Size == 4 || Op->Size == 8, "Invalid size");
LOGMAN_THROW_A_FMT(Op->Size == IR::OpSize::i32Bit || Op->Size == IR::OpSize::i64Bit, "Invalid size");
return ConvertSize(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(uint8_t ElementSize) {
LOGMAN_THROW_AA_FMT(ElementSize == 1 || ElementSize == 2 || ElementSize == 4 || ElementSize == 8 || ElementSize == 16, "Invalid size");
return ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ARMEmitter::SubRegSize::i128Bit;
ARMEmitter::SubRegSize ConvertSubRegSize16(IR::OpSize ElementSize) {
LOGMAN_THROW_A_FMT(ElementSize == IR::OpSize::i8Bit || ElementSize == IR::OpSize::i16Bit || ElementSize == IR::OpSize::i32Bit ||
ElementSize == IR::OpSize::i64Bit || ElementSize == IR::OpSize::i128Bit,
"Invalid size");
return ElementSize == IR::OpSize::i8Bit ? ARMEmitter::SubRegSize::i8Bit :
ElementSize == IR::OpSize::i16Bit ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == IR::OpSize::i32Bit ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == IR::OpSize::i64Bit ? ARMEmitter::SubRegSize::i64Bit :
ARMEmitter::SubRegSize::i128Bit;
}
[[nodiscard]]
@@ -154,8 +157,8 @@ private:
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(uint8_t ElementSize) {
LOGMAN_THROW_AA_FMT(ElementSize != 16, "Invalid size");
ARMEmitter::SubRegSize ConvertSubRegSize8(IR::OpSize ElementSize) {
LOGMAN_THROW_A_FMT(ElementSize != IR::OpSize::i128Bit, "Invalid size");
return ConvertSubRegSize16(ElementSize);
}
@@ -166,13 +169,13 @@ private:
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 8, "Invalid size");
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i64Bit, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i8Bit, "Invalid size");
return ConvertSubRegSize8(Op);
}
@@ -183,13 +186,13 @@ private:
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 16, "Invalid size");
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i128Bit, "Invalid size");
return ConvertSubRegSizePair16(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i8Bit, "Invalid size");
return ConvertSubRegSizePair8(Op);
}
@@ -226,16 +229,20 @@ private:
bool IsGPR(IR::NodeID Node) const;
[[nodiscard]]
ARMEmitter::ExtendedMemOperand GenerateMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
ARMEmitter::ExtendedMemOperand GenerateMemOperand(IR::OpSize AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
[[nodiscard]]
ARMEmitter::Register ApplyMemOperand(IR::OpSize AccessSize, ARMEmitter::Register Base, ARMEmitter::Register Tmp,
IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
// NOTE: Will use TMP1 as a way to encode immediates that happen to fall outside
// the limits of the scalar plus immediate variant of SVE load/stores.
//
// TMP1 is safe to use again once this memory operand is used with its
// equivalent loads or stores that this was called for.
[[nodiscard]]
ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
ARMEmitter::SVEMemOperand GenerateSVEMemOperand(IR::OpSize AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
[[nodiscard]]
@@ -316,20 +323,24 @@ private:
using ScalarFMAOpCaller =
std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2, ARMEmitter::VRegister Src3)>;
void VFScalarFMAOperation(uint8_t OpSize, uint8_t ElementSize, ScalarFMAOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
void VFScalarFMAOperation(IR::OpSize OpSize, IR::OpSize ElementSize, ScalarFMAOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Upper, ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2,
ARMEmitter::VRegister Addend);
using ScalarBinaryOpCaller = std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2)>;
void VFScalarOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarBinaryOpCaller ScalarEmit,
void VFScalarOperation(IR::OpSize OpSize, IR::OpSize ElementSize, bool ZeroUpperBits, ScalarBinaryOpCaller ScalarEmit,
ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2);
using ScalarUnaryOpCaller = std::function<void(ARMEmitter::VRegister Dst, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> SrcVar)>;
void VFScalarUnaryOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarUnaryOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Vector1, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> Vector2);
void VFScalarUnaryOperation(IR::OpSize OpSize, IR::OpSize ElementSize, bool ZeroUpperBits, ScalarUnaryOpCaller ScalarEmit,
ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1,
std::variant<ARMEmitter::VRegister, ARMEmitter::Register> Vector2);
void Emulate128BitGather(size_t Size, size_t ElementSize, ARMEmitter::VRegister Dst, ARMEmitter::VRegister IncomingDst,
void Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize, ARMEmitter::VRegister Dst, ARMEmitter::VRegister IncomingDst,
std::optional<ARMEmitter::Register> BaseAddr, ARMEmitter::VRegister VectorIndexLow,
std::optional<ARMEmitter::VRegister> VectorIndexHigh, ARMEmitter::VRegister MaskReg, size_t VectorIndexSize,
std::optional<ARMEmitter::VRegister> VectorIndexHigh, ARMEmitter::VRegister MaskReg, IR::OpSize VectorIndexSize,
size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale);
void EmitInterruptChecks(bool CheckTF);
// Runtime selection;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
@@ -352,4 +363,7 @@ private:
#undef DEF_OP
};
[[nodiscard]]
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
} // namespace FEXCore::CPU
@@ -1,21 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/Core/CPUBackend.h"
#include <FEXCore/fextl/memory.h>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]]
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
} // namespace FEXCore::CPU
@@ -10,7 +10,7 @@ $end_info$
#endif
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include <FEXCore/Core/SignalDelegator.h>
@@ -148,7 +148,7 @@ DEF_OP(PushRoundingMode) {
} else if (Op->RoundMode == 0) {
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(3 << 22));
} else {
LOGMAN_THROW_AA_FMT(Op->RoundMode == 1 || Op->RoundMode == 2, "expect a valid round mode");
LOGMAN_THROW_A_FMT(Op->RoundMode == 1 || Op->RoundMode == 2, "expect a valid round mode");
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(Op->RoundMode << 22));
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, (Op->RoundMode == 2 ? 1 : 2) << 22);
@@ -192,8 +192,17 @@ DEF_OP(Print) {
PopDynamicRegsAndLR();
}
#ifndef _WIN32
DEF_OP(ProcessorID) {
if (CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
mrs(GetReg(Node), ARMEmitter::SystemRegister::TPIDRRO_EL0);
return;
}
#ifdef _WIN32
else {
// If on Windows and TPIDRRO isn't supported (like in wine), then this is a programming error.
ERROR_AND_DIE_FMT("Unsupported");
}
#else
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
@@ -248,12 +257,8 @@ DEF_OP(ProcessorID) {
// CPU is in w0
// Node is in w1
orr(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0, ARMEmitter::Reg::r1, ARMEmitter::ShiftType::LSL, 12);
}
#else
DEF_OP(ProcessorID) {
ERROR_AND_DIE_FMT("Unsupported");
}
#endif
}
DEF_OP(RDRAND) {
auto Op = IROp->C<IR::IROp_RDRAND>();
@@ -5,7 +5,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/JIT/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
@@ -44,11 +44,11 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = PagePointer + ctx->Config.VirtualMemSize / 4096 * 8;
LOGMAN_THROW_AA_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = PageMemory + CODE_SIZE;
LOGMAN_THROW_AA_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
}
+1 -1
View File
@@ -90,7 +90,7 @@ public:
std::lock_guard<std::recursive_mutex> lk(WriteLock);
[[maybe_unused]] auto Inserted = BlockList.emplace(Address, (uintptr_t)HostCode).second;
LOGMAN_THROW_AA_FMT(Inserted, "Duplicate block mapping added");
LOGMAN_THROW_A_FMT(Inserted, "Duplicate block mapping added");
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -43,65 +43,72 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref NewVec = _VExtr(16, 8, Dest, Src, 1);
Ref NewVec = _VExtr(OpSize::i128Bit, OpSize::i64Bit, Dest, Src, 1);
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
Ref Result = _VXor(16, 1, Dest, NewVec);
Ref Result = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, NewVec);
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
// This instruction mostly matches ARMv8's SHA1SU1 instruction but one of the elements are flipped in an unexpected way.
// Do all the work without it.
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
// ARM SHA1 mostly matches x86 semantics, except the input and outputs are both flipped from elements 0,1,2,3 to 3,2,1,0.
auto Src1 = SHADataShuffle(Dest);
auto Src2 = SHADataShuffle(Src);
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
// The result is swizzled differently than expected
Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
} else {
// Shift the incoming source left by a 32-bit element, inserting Zeros.
// This could be slightly improved to use a VInsGPR with the zero register.
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
auto Src2Shift = _VExtr(OpSize::i128Bit, OpSize::i8Bit, Src, ZeroRegister, 12);
auto Xor1 = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, Src2Shift);
// Shift the incoming source left by a 32-bit element, inserting Zeros.
// This could be slightly improved to use a VInsGPR with the zero register.
auto Src2Shift = _VExtr(OpSize::i128Bit, OpSize::i8Bit, Src, ZeroRegister, 12);
auto Xor1 = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, Src2Shift);
// Emulate rotate.
auto ShiftLeftXor1 = _VShlI(OpSize::i128Bit, OpSize::i32Bit, Xor1, 1);
auto RotatedXor1 = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXor1, Xor1, 31);
// Emulate rotate.
auto ShiftLeftXor1 = _VShlI(OpSize::i128Bit, OpSize::i32Bit, Xor1, 1);
auto RotatedXor1 = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXor1, Xor1, 31);
// Element0 didn't get XOR'd with anything, so do it now.
auto ExtractUpper = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, RotatedXor1, 3);
auto XorLower = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, ExtractUpper);
// Element0 didn't get XOR'd with anything, so do it now.
auto ExtractUpper = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, RotatedXor1, 3);
auto XorLower = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, ExtractUpper);
// Emulate rotate.
auto ShiftLeftXorLower = _VShlI(OpSize::i128Bit, OpSize::i32Bit, XorLower, 1);
auto RotatedXorLower = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXorLower, XorLower, 31);
// Emulate rotate.
auto ShiftLeftXorLower = _VShlI(OpSize::i128Bit, OpSize::i32Bit, XorLower, 1);
auto RotatedXorLower = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXorLower, XorLower, 31);
Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 0, 0, RotatedXor1, RotatedXorLower);
}
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 0, 0, RotatedXor1, RotatedXorLower);
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
using FnType = Ref (*)(OpDispatchBuilder&, Ref, Ref, Ref);
const auto f0 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
const auto f0 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1c?
return Self._Xor(OpSize::i32Bit, Self._And(OpSize::i32Bit, B, C), Self._Andn(OpSize::i32Bit, D, B));
};
const auto f1 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
const auto f1 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1p with different key
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
const auto f2 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
const auto f2 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1m
return Self.BitwiseAtLeastTwo(B, C, D);
};
const auto f3 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
const auto f3 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1p
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
@@ -119,58 +126,92 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
f3,
};
const uint64_t Imm8 = Op->Src[1].Literal() & 0b11;
const FnType Fn = fn_array[Imm8];
auto K = _Constant(32, k_array[Imm8]);
const uint64_t Imm8 = Op->Src[1].Literal() & 0b11;
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
auto W0E = _VExtractToGPR(16, 4, Src, 3);
Ref Result {};
if (CTX->HostFeatures.SupportsSHA) {
Ref ConstantVector {};
switch (Imm8) {
case 0:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K0);
break;
case 1:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K1);
break;
case 2:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K2);
break;
case 3:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K3);
break;
}
using RoundResult = std::tuple<Ref, Ref, Ref, Ref, Ref>;
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
const auto Round0 = [&]() -> RoundResult {
auto A = _VExtractToGPR(16, 4, Dest, 3);
auto B = _VExtractToGPR(16, 4, Dest, 2);
auto C = _VExtractToGPR(16, 4, Dest, 1);
auto D = _VExtractToGPR(16, 4, Dest, 0);
Ref Src1 = SHADataShuffle(Dest);
Ref Src2 = SHADataShuffle(Src);
Src2 = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src2, ConstantVector);
auto A1 =
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), W0E), K);
auto B1 = A;
auto C1 = _Ror(OpSize::i32Bit, B, _Constant(32, 2));
auto D1 = C;
auto E1 = D;
switch (Imm8) {
case 0: Result = SHADataShuffle(_VSha1C(Src1, ZeroRegister, Src2)); break;
case 2: Result = SHADataShuffle(_VSha1M(Src1, ZeroRegister, Src2)); break;
case 1:
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
} else {
const FnType Fn = fn_array[Imm8];
auto K = _Constant(OpSize::i32Bit, k_array[Imm8]);
auto W0E = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](Ref A, Ref B, Ref C, Ref D, Ref E, Ref Src, unsigned W_idx) -> RoundResult {
// Kill W and E at the beginning
auto W = _VExtractToGPR(16, 4, Src, W_idx);
auto Q = _Add(OpSize::i32Bit, W, E);
using RoundResult = std::tuple<Ref, Ref, Ref, Ref, Ref>;
auto ANext =
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), Q), K);
auto BNext = A;
auto CNext = _Ror(OpSize::i32Bit, B, _Constant(32, 2));
auto DNext = C;
auto ENext = D;
const auto Round0 = [&]() -> RoundResult {
auto A = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto B = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto C = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
auto D = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
return {ANext, BNext, CNext, DNext, ENext};
};
auto A1 =
_Add(OpSize::i32Bit,
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 27))), W0E), K);
auto B1 = A;
auto C1 = _Ror(OpSize::i32Bit, B, _Constant(OpSize::i32Bit, 2));
auto D1 = C;
auto E1 = D;
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, Src, 2);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, Src, 1);
auto Final = Round1To3(A3, B3, C3, D3, E3, Src, 0);
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](Ref A, Ref B, Ref C, Ref D, Ref E, Ref Src, unsigned W_idx) -> RoundResult {
// Kill W and E at the beginning
auto W = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, W_idx);
auto Q = _Add(OpSize::i32Bit, W, E);
auto Dest3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(16, 4, 2, Dest3, std::get<1>(Final));
auto Dest1 = _VInsGPR(16, 4, 1, Dest2, std::get<2>(Final));
auto Dest0 = _VInsGPR(16, 4, 0, Dest1, std::get<3>(Final));
auto ANext =
_Add(OpSize::i32Bit,
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 27))), Q), K);
auto BNext = A;
auto CNext = _Ror(OpSize::i32Bit, B, _Constant(OpSize::i32Bit, 2));
auto DNext = C;
auto ENext = D;
StoreResult(FPRClass, Op, Dest0, -1);
return {ANext, BNext, CNext, DNext, ENext};
};
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, Src, 2);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, Src, 1);
auto Final = Round1To3(A3, B3, C3, D3, E3, Src, 0);
auto Dest3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, Dest3, std::get<1>(Final));
auto Dest1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, Dest2, std::get<2>(Final));
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, Dest1, std::get<3>(Final));
}
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
@@ -183,52 +224,65 @@ void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
Result = _VSha256U0(Dest, Src);
} else {
const auto Sigma0 = [this](Ref W) -> Ref {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 7)), _Ror(OpSize::i32Bit, W, _Constant(32, 18))),
_Lshr(OpSize::i32Bit, W, _Constant(32, 3)));
return _Xor(
OpSize::i32Bit,
_Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 7)), _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 18))),
_Lshr(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 3)));
};
auto W4 = _VExtractToGPR(16, 4, Src, 0);
auto W3 = _VExtractToGPR(16, 4, Dest, 3);
auto W2 = _VExtractToGPR(16, 4, Dest, 2);
auto W1 = _VExtractToGPR(16, 4, Dest, 1);
auto W0 = _VExtractToGPR(16, 4, Dest, 0);
auto W4 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 0);
auto W3 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto W2 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto W1 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
auto W0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
auto Sig3 = _Add(OpSize::i32Bit, W3, Sigma0(W4));
auto Sig2 = _Add(OpSize::i32Bit, W2, Sigma0(W3));
auto Sig1 = _Add(OpSize::i32Bit, W1, Sigma0(W2));
auto Sig0 = _Add(OpSize::i32Bit, W0, Sigma0(W1));
auto D3 = _VInsGPR(16, 4, 3, Dest, Sig3);
auto D2 = _VInsGPR(16, 4, 2, D3, Sig2);
auto D1 = _VInsGPR(16, 4, 1, D2, Sig1);
Result = _VInsGPR(16, 4, 0, D1, Sig0);
auto D3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, Sig3);
auto D2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, D3, Sig2);
auto D1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, D2, Sig1);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, D1, Sig0);
}
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
const auto Sigma1 = [this](Ref W) -> Ref {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 17)), _Ror(OpSize::i32Bit, W, _Constant(32, 19))),
_Lshr(OpSize::i32Bit, W, _Constant(32, 10)));
return _Xor(
OpSize::i32Bit,
_Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 17)), _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 19))),
_Lshr(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 10)));
};
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
auto W14 = _VExtractToGPR(16, 4, Src, 2);
auto W15 = _VExtractToGPR(16, 4, Src, 3);
auto W16 = _Add(OpSize::i32Bit, _VExtractToGPR(16, 4, Dest, 0), Sigma1(W14));
auto W17 = _Add(OpSize::i32Bit, _VExtractToGPR(16, 4, Dest, 1), Sigma1(W15));
auto W18 = _Add(OpSize::i32Bit, _VExtractToGPR(16, 4, Dest, 2), Sigma1(W16));
auto W19 = _Add(OpSize::i32Bit, _VExtractToGPR(16, 4, Dest, 3), Sigma1(W17));
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
auto Src1 = _VExtr(OpSize::i128Bit, OpSize::i32Bit, Dest, Dest, 3);
auto DupDst = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Src2 = _VZip2(OpSize::i128Bit, OpSize::i64Bit, DupDst, Src);
auto D3 = _VInsGPR(16, 4, 3, Dest, W19);
auto D2 = _VInsGPR(16, 4, 2, D3, W18);
auto D1 = _VInsGPR(16, 4, 1, D2, W17);
auto D0 = _VInsGPR(16, 4, 0, D1, W16);
Result = _VSha256U1(Src1, Src2);
} else {
auto W14 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 2);
auto W15 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
auto W16 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0), Sigma1(W14));
auto W17 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1), Sigma1(W15));
auto W18 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2), Sigma1(W16));
auto W19 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3), Sigma1(W17));
StoreResult(FPRClass, Op, D0, -1);
auto D3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, W19);
auto D2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, D3, W18);
auto D1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, D2, W17);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, D1, W16);
}
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
Ref OpDispatchBuilder::BitwiseAtLeastTwo(Ref A, Ref B, Ref C) {
@@ -246,12 +300,12 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
return _Xor(OpSize::i32Bit, _And(OpSize::i32Bit, E, F), _Andn(OpSize::i32Bit, G, E));
};
const auto Sigma0 = [this](Ref A) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(32, 2)), A, ShiftType::ROR, 13), A,
ShiftType::ROR, 22);
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 2)), A, ShiftType::ROR, 13),
A, ShiftType::ROR, 22);
};
const auto Sigma1 = [this](Ref E) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(32, 6)), E, ShiftType::ROR, 11), E,
ShiftType::ROR, 25);
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(OpSize::i32Bit, 6)), E, ShiftType::ROR, 11),
E, ShiftType::ROR, 25);
};
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
@@ -259,64 +313,64 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
auto E0 = _VExtractToGPR(16, 4, Src, 1);
auto F0 = _VExtractToGPR(16, 4, Src, 0);
auto G0 = _VExtractToGPR(16, 4, Dest, 1);
auto E0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 1);
auto F0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 0);
auto G0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
Ref Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
auto WK0 = _VExtractToGPR(16, 4, XMM0, 0);
auto WK0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, XMM0, 0);
Q0 = _Add(OpSize::i32Bit, Q0, WK0);
auto H0 = _VExtractToGPR(16, 4, Dest, 0);
auto H0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
Q0 = _Add(OpSize::i32Bit, Q0, H0);
auto A0 = _VExtractToGPR(16, 4, Src, 3);
auto B0 = _VExtractToGPR(16, 4, Src, 2);
auto C0 = _VExtractToGPR(16, 4, Dest, 3);
auto A0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
auto B0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 2);
auto C0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto A1 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q0, BitwiseAtLeastTwo(A0, B0, C0)), Sigma0(A0));
auto D0 = _VExtractToGPR(16, 4, Dest, 2);
auto D0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto E1 = _Add(OpSize::i32Bit, Q0, D0);
Ref Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
auto WK1 = _VExtractToGPR(16, 4, XMM0, 1);
auto WK1 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, XMM0, 1);
Q1 = _Add(OpSize::i32Bit, Q1, WK1);
// Rematerialize G0. Costs a move but saves spilling, coming out ahead.
G0 = _VExtractToGPR(16, 4, Dest, 1);
G0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
Q1 = _Add(OpSize::i32Bit, Q1, G0);
auto A2 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q1, BitwiseAtLeastTwo(A1, A0, B0)), Sigma0(A1));
// Rematerialize C0. As with G0.
C0 = _VExtractToGPR(16, 4, Dest, 3);
C0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto E2 = _Add(OpSize::i32Bit, Q1, C0);
auto Res3 = _VInsGPR(16, 4, 3, Dest, A2);
auto Res2 = _VInsGPR(16, 4, 2, Res3, A1);
auto Res1 = _VInsGPR(16, 4, 1, Res2, E2);
auto Res0 = _VInsGPR(16, 4, 0, Res1, E1);
auto Res3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, A2);
auto Res2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, Res3, A1);
auto Res1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, Res2, E2);
auto Res0 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, Res1, E1);
StoreResult(FPRClass, Op, Res0, -1);
StoreResult(FPRClass, Op, Res0, OpSize::iInvalid);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESImc(Src);
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEnc(16, Dest, Src, LoadZeroVector(16));
StoreResult(FPRClass, Op, Result, -1);
Ref Result = _VAESEnc(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
@@ -325,19 +379,19 @@ void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEnc(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEncLast(16, Dest, Src, LoadZeroVector(16));
StoreResult(FPRClass, Op, Result, -1);
Ref Result = _VAESEncLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
@@ -346,19 +400,19 @@ void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEncLast(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDec(16, Dest, Src, LoadZeroVector(16));
StoreResult(FPRClass, Op, Result, -1);
Ref Result = _VAESDec(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
@@ -367,19 +421,19 @@ void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDec(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDecLast(16, Dest, Src, LoadZeroVector(16));
StoreResult(FPRClass, Op, Result, -1);
Ref Result = _VAESDecLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
@@ -388,20 +442,20 @@ void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDecLast(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const uint64_t RCON = Op->Src[1].Literal();
auto KeyGenSwizzle = LoadAndCacheNamedVectorConstant(16, NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE);
return _VAESKeyGenAssist(Src, KeyGenSwizzle, LoadZeroVector(16), RCON);
auto KeyGenSwizzle = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE);
return _VAESKeyGenAssist(Src, KeyGenSwizzle, LoadZeroVector(OpSize::i128Bit), RCON);
}
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
Ref Result = AESKeyGenAssistImpl(Op);
StoreResult(FPRClass, Op, Result, -1);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
@@ -409,19 +463,19 @@ void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Literal());
auto Res = _PCLMUL(16, Dest, Src, Selector & 0b1'0001);
StoreResult(FPRClass, Op, Res, -1);
auto Res = _PCLMUL(OpSize::i128Bit, Dest, Src, Selector & 0b1'0001);
StoreResult(FPRClass, Op, Res, OpSize::iInvalid);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto DstSize = OpSizeFromDst(Op);
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Literal());
Ref Res = _PCLMUL(DstSize, Src1, Src2, Selector & 0b1'0001);
StoreResult(FPRClass, Op, Res, -1);
StoreResult(FPRClass, Op, Res, OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -5,41 +5,41 @@
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_DDDTable[] = {
{0x0C, 1, &OpDispatchBuilder::PI2FWOp},
{0x0D, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<4, false>},
{0x0D, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x1C, 1, &OpDispatchBuilder::PF2IWOp},
{0x1D, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, false>},
{0x1D, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x86, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECP, 4>},
{0x87, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRSQRT, 4>},
{0x86, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x87, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, false>},
{0x8A, 1, &OpDispatchBuilder::PFNACCOp},
{0x8E, 1, &OpDispatchBuilder::PFPNACCOp},
{0x90, 1, &OpDispatchBuilder::VPFCMPOp<1>},
{0x94, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, 4>},
{0x96, 1, &OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRECP, 4>},
{0x97, 1, &OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRSQRT, 4>},
{0x94, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x96, 1, &OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x97, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, true>},
{0x9A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, 4>},
{0x9E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, 4>},
{0x9A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x9E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0xA0, 1, &OpDispatchBuilder::VPFCMPOp<2>},
{0xA4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, 4>},
{0xA4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
// Can be treated as a move
{0xA6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0xA7, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0xAA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VFSUB, 4>},
{0xAE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, 4>},
{0xAA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VFSUB, OpSize::i32Bit>},
{0xAE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0xB0, 1, &OpDispatchBuilder::VPFCMPOp<0>},
{0xB4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, 4>},
{0xB4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
// Can be treated as a move
{0xB6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0xB7, 1, &OpDispatchBuilder::PMULHRWOp},
{0xBB, 1, &OpDispatchBuilder::PSWAPDOp},
{0xBF, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, 1>},
{0xBF, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i8Bit>},
};
} // namespace FEXCore::IR
@@ -19,7 +19,7 @@ $end_info$
namespace FEXCore::IR {
constexpr std::array<uint32_t, 17> FlagOffsets = {
FEXCore::X86State::RFLAG_CF_RAW_LOC, FEXCore::X86State::RFLAG_PF_RAW_LOC, FEXCore::X86State::RFLAG_AF_RAW_LOC,
FEXCore::X86State::RFLAG_ZF_RAW_LOC, FEXCore::X86State::RFLAG_SF_RAW_LOC, FEXCore::X86State::RFLAG_TF_LOC,
FEXCore::X86State::RFLAG_ZF_RAW_LOC, FEXCore::X86State::RFLAG_SF_RAW_LOC, FEXCore::X86State::RFLAG_TF_RAW_LOC,
FEXCore::X86State::RFLAG_IF_LOC, FEXCore::X86State::RFLAG_DF_RAW_LOC, FEXCore::X86State::RFLAG_OF_RAW_LOC,
FEXCore::X86State::RFLAG_IOPL_LOC, FEXCore::X86State::RFLAG_NT_LOC, FEXCore::X86State::RFLAG_RF_LOC,
FEXCore::X86State::RFLAG_VM_LOC, FEXCore::X86State::RFLAG_AC_LOC, FEXCore::X86State::RFLAG_VIF_LOC,
@@ -36,13 +36,9 @@ void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, Ref Src) {
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
// Calculate flags early.
// Could use InvalidateDeferredFlags() if we had masked invalidation.
// This is only a partial overwrite of flags since OF isn't stored here.
CalculateDeferredFlags();
NumFlags = 5;
} else {
// We are overwriting all RFLAGS. Invalidate the deferred flag state.
InvalidateDeferredFlags();
}
// PF and CF are both stored inverted, so hoist the invert.
@@ -138,9 +134,9 @@ Ref OpDispatchBuilder::GetPackedRFLAG(uint32_t FlagsMask) {
return Original;
}
void OpDispatchBuilder::CalculateOF(uint8_t SrcSize, Ref Res, Ref Src1, Ref Src2, bool Sub) {
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
uint64_t SignBit = (SrcSize * 8) - 1;
void OpDispatchBuilder::CalculateOF(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2, bool Sub) {
const auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
const uint64_t SignBit = IR::OpSizeAsBits(SrcSize) - 1;
Ref Anded = nullptr;
// For add, OF is set iff the sources have the same sign but the destination
@@ -171,7 +167,7 @@ void OpDispatchBuilder::CalculateOF(uint8_t SrcSize, Ref Res, Ref Src1, Ref Src2
}
}
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Anded, SrcSize * 8 - 1, true);
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Anded, SignBit, true);
}
Ref OpDispatchBuilder::LoadPFRaw(bool Mask, bool Invert) {
@@ -189,8 +185,9 @@ Ref OpDispatchBuilder::LoadAF() {
// Read the result, stored for PF.
auto Result = GetRFLAG(FEXCore::X86State::RFLAG_PF_RAW_LOC);
// What's left is to XOR and extract. This is the deferred part.
return _Bfe(OpSize::i32Bit, 1, 4, _Xor(OpSize::i32Bit, AFWord, Result));
// What's left is to XOR and extract. This is the deferred part. We
// specifically use a 64-bit Xor here as we don't need masking.
return _Bfe(OpSize::i32Bit, 1, 4, _Xor(OpSize::i64Bit, AFWord, Result));
}
void OpDispatchBuilder::FixupAF() {
@@ -203,7 +200,8 @@ void OpDispatchBuilder::FixupAF() {
auto PFRaw = GetRFLAG(FEXCore::X86State::RFLAG_PF_RAW_LOC);
auto AFRaw = GetRFLAG(FEXCore::X86State::RFLAG_AF_RAW_LOC);
Ref XorRes = _Xor(OpSize::i32Bit, AFRaw, PFRaw);
// Again 64-bit as masking is more expensive given our ConstProp design.
Ref XorRes = _Xor(OpSize::i64Bit, AFRaw, PFRaw);
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
@@ -242,8 +240,8 @@ void OpDispatchBuilder::CalculateAF(Ref Src1, Ref Src2) {
// We store the XOR of the arguments. At read time, we XOR with the
// appropriate bit of the result (available as the PF flag) and extract the
// appropriate bit.
Ref XorRes = _Xor(OpSize::i32Bit, Src1, Src2);
// appropriate bit. Again 64-bit to avoid masking.
Ref XorRes = _Xor(OpSize::i64Bit, Src1, Src2);
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
@@ -262,22 +260,22 @@ Ref OpDispatchBuilder::IncrementByCarry(OpSize OpSize, Ref Src) {
return _NZCVSelectIncrement(OpSize, {CFInverted ? COND_UGE : COND_ULT}, Src, Src);
}
Ref OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, Ref Src1, Ref Src2) {
Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2) {
auto Zero = _InlineConstant(0);
auto One = _InlineConstant(1);
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
Ref Res;
CalculateAF(Src1, Src2);
if (SrcSize >= 4) {
if (SrcSize >= OpSize::i32Bit) {
RectifyCarryInvert(false);
HandleNZCV_RMW();
Res = _AdcWithFlags(OpSize, Src1, Src2);
CFInverted = false;
} else {
// Need to zero-extend for correct comparisons below
Src2 = _Bfe(OpSize, SrcSize * 8, 0, Src2);
Src2 = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Src2);
// Note that we do not extend Src2PlusCF, since we depend on proper
// 32-bit arithmetic to correctly handle the Src2 = 0xffff case.
@@ -285,7 +283,7 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, Ref Src1, Ref Src2) {
// Need to zero-extend for the comparison.
Res = _Add(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, SrcSize * 8, 0, Res);
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
// TODO: We can fold that second Bfe in (cmp uxth).
auto SelectCFInv = _Select(FEXCore::IR::COND_UGE, Res, Src2PlusCF, One, Zero);
@@ -299,15 +297,15 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, Ref Src1, Ref Src2) {
return Res;
}
Ref OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, Ref Src1, Ref Src2) {
Ref OpDispatchBuilder::CalculateFlags_SBB(IR::OpSize SrcSize, Ref Src1, Ref Src2) {
auto Zero = _InlineConstant(0);
auto One = _InlineConstant(1);
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
CalculateAF(Src1, Src2);
Ref Res;
if (SrcSize >= 4) {
if (SrcSize >= OpSize::i32Bit) {
// Arm's subtraction has inverted CF from x86, so rectify the input and
// invert the output.
RectifyCarryInvert(true);
@@ -316,13 +314,13 @@ Ref OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, Ref Src1, Ref Src2) {
CFInverted = true;
} else {
// Zero extend for correct comparison behaviour with Src1 = 0xffff.
Src1 = _Bfe(OpSize, SrcSize * 8, 0, Src1);
Src2 = _Bfe(OpSize, SrcSize * 8, 0, Src2);
Src1 = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Src1);
Src2 = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Src2);
auto Src2PlusCF = IncrementByCarry(OpSize, Src2);
Res = _Sub(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, SrcSize * 8, 0, Res);
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
auto SelectCFInv = _Select(FEXCore::IR::COND_UGE, Src1, Src2PlusCF, One, Zero);
@@ -335,7 +333,7 @@ Ref OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, Ref Src1, Ref Src2) {
return Res;
}
Ref OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, Ref Src1, Ref Src2, bool UpdateCF) {
Ref OpDispatchBuilder::CalculateFlags_SUB(IR::OpSize SrcSize, Ref Src1, Ref Src2, bool UpdateCF) {
// Stash CF before stomping over it
auto OldCFInv = UpdateCF ? nullptr : GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC, true);
@@ -344,10 +342,10 @@ Ref OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, Ref Src1, Ref Src2, b
CalculateAF(Src1, Src2);
Ref Res;
if (SrcSize >= 4) {
Res = _SubWithFlags(IR::SizeToOpSize(SrcSize), Src1, Src2);
if (SrcSize >= OpSize::i32Bit) {
Res = _SubWithFlags(SrcSize, Src1, Src2);
} else {
_SubNZCV(IR::SizeToOpSize(SrcSize), Src1, Src2);
_SubNZCV(SrcSize, Src1, Src2);
Res = _Sub(OpSize::i32Bit, Src1, Src2);
}
@@ -365,7 +363,7 @@ Ref OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, Ref Src1, Ref Src2, b
return Res;
}
Ref OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, Ref Src1, Ref Src2, bool UpdateCF) {
Ref OpDispatchBuilder::CalculateFlags_ADD(IR::OpSize SrcSize, Ref Src1, Ref Src2, bool UpdateCF) {
// Stash CF before stomping over it
auto OldCFInv = UpdateCF ? nullptr : GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC, true);
@@ -374,10 +372,10 @@ Ref OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, Ref Src1, Ref Src2, b
CalculateAF(Src1, Src2);
Ref Res;
if (SrcSize >= 4) {
Res = _AddWithFlags(IR::SizeToOpSize(SrcSize), Src1, Src2);
if (SrcSize >= OpSize::i32Bit) {
Res = _AddWithFlags(SrcSize, Src1, Src2);
} else {
_AddNZCV(IR::SizeToOpSize(SrcSize), Src1, Src2);
_AddNZCV(SrcSize, Src1, Src2);
Res = _Add(OpSize::i32Bit, Src1, Src2);
}
@@ -394,13 +392,13 @@ Ref OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, Ref Src1, Ref Src2, b
return Res;
}
void OpDispatchBuilder::CalculateFlags_MUL(uint8_t SrcSize, Ref Res, Ref High) {
void OpDispatchBuilder::CalculateFlags_MUL(IR::OpSize SrcSize, Ref Res, Ref High) {
HandleNZCVWrite();
InvalidatePF_AF();
// CF and OF are set if the result of the operation can't be fit in to the destination register
// If the value can fit then the top bits will be zero
auto SignBit = _Sbfe(OpSize::i64Bit, 1, SrcSize * 8 - 1, Res);
auto SignBit = _Sbfe(OpSize::i64Bit, 1, IR::OpSizeAsBits(SrcSize) - 1, Res);
_SubNZCV(OpSize::i64Bit, High, SignBit);
// If High = SignBit, then sets to nZCv. Else sets to nzcV. Since SF/ZF
@@ -415,7 +413,7 @@ void OpDispatchBuilder::CalculateFlags_UMUL(Ref High) {
InvalidatePF_AF();
auto Zero = _InlineConstant(0);
OpSize Size = IR::SizeToOpSize(GetOpSize(High));
const auto Size = GetOpSize(High);
// CF and OF are set if the result of the operation can't be fit in to the destination register
// The result register will be all zero if it can't fit due to how multiplication behaves
@@ -427,7 +425,7 @@ void OpDispatchBuilder::CalculateFlags_UMUL(Ref High) {
CFInverted = true;
}
void OpDispatchBuilder::CalculateFlags_Logical(uint8_t SrcSize, Ref Res, Ref Src1, Ref Src2) {
void OpDispatchBuilder::CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2) {
InvalidateAF();
CalculatePF(Res);
@@ -436,13 +434,13 @@ void OpDispatchBuilder::CalculateFlags_Logical(uint8_t SrcSize, Ref Res, Ref Src
SetNZ_ZeroCV(SrcSize, Res);
}
void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, Ref UnmaskedRes, Ref Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref UnmaskedRes, Ref Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) {
return;
}
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
SetNZ_ZeroCV(SrcSize, UnmaskedRes);
@@ -451,7 +449,7 @@ void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, Ref U
// Extract the last bit shifted in to CF. Shift is already masked, but for
// 8/16-bit it might be >= SrcSizeBits, in which case CF is cleared. There's
// nothing to do in that case since we already cleared CF above.
auto SrcSizeBits = SrcSize * 8;
const auto SrcSizeBits = IR::OpSizeAsBits(SrcSize);
if (Shift < SrcSizeBits) {
SetCFDirect(Src1, SrcSizeBits - Shift, true);
}
@@ -464,13 +462,13 @@ void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, Ref U
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
if (Shift == 1) {
auto Xor = _Xor(OpSize, UnmaskedRes, Src1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Xor, SrcSize * 8 - 1, true);
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Xor, IR::OpSizeAsBits(SrcSize) - 1, true);
} else {
// Undefined, we choose to zero as part of SetNZ_ZeroCV
}
}
void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) {
return;
@@ -490,7 +488,7 @@ void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(uint8_t SrcSize,
// already zeroed there's nothing to do here.
}
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
// Set SF and PF. Clobbers OF, but OF only defined for Shift = 1 where it is
// set below.
SetNZ_ZeroCV(SrcSize, Res);
@@ -502,7 +500,7 @@ void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(uint8_t SrcSize
InvalidateAF();
}
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediate(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) {
return;
@@ -515,18 +513,18 @@ void OpDispatchBuilder::CalculateFlags_ShiftRightImmediate(uint8_t SrcSize, Ref
// Only defined when Shift is 1 else undefined
// Is set to the MSB of the original value
if (Shift == 1) {
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Src1, SrcSize * 8 - 1, true);
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Src1, IR::OpSizeAsBits(SrcSize) - 1, true);
}
}
}
void OpDispatchBuilder::CalculateFlags_ShiftRightDoubleImmediate(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculateFlags_ShiftRightDoubleImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) {
return;
}
const auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
const auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
CalculateFlags_ShiftRightImmediateCommon(SrcSize, Res, Src1, Shift);
// OF
@@ -536,12 +534,12 @@ void OpDispatchBuilder::CalculateFlags_ShiftRightDoubleImmediate(uint8_t SrcSize
// XOR of Result and Src1
if (Shift == 1) {
auto val = _Xor(OpSize, Src1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(val, SrcSize * 8 - 1, true);
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(val, IR::OpSizeAsBits(SrcSize) - 1, true);
}
}
}
void OpDispatchBuilder::CalculateFlags_ZCNT(uint8_t SrcSize, Ref Result) {
void OpDispatchBuilder::CalculateFlags_ZCNT(IR::OpSize SrcSize, Ref Result) {
// OF, SF, AF, PF all undefined
// Test ZF of result, SF is undefined so this is ok.
SetNZ_ZeroCV(SrcSize, Result);
@@ -549,7 +547,7 @@ void OpDispatchBuilder::CalculateFlags_ZCNT(uint8_t SrcSize, Ref Result) {
// Now set CF if the Result = SrcSize * 8. Since SrcSize is a power-of-two and
// Result is <= SrcSize * 8, we equivalently check if the log2(SrcSize * 8)
// bit is set. No masking is needed because no higher bits could be set.
unsigned CarryBit = FEXCore::ilog2(SrcSize * 8u);
unsigned CarryBit = FEXCore::ilog2(IR::OpSizeAsBits(SrcSize));
SetCFDirect(Result, CarryBit);
}
@@ -11,64 +11,64 @@ constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_66, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_NONE, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, 2>},
{OPD(PF_38_66, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, 2>},
{OPD(PF_38_NONE, 0x02), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, 4>},
{OPD(PF_38_66, 0x02), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, 4>},
{OPD(PF_38_NONE, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, OpSize::i16Bit>},
{OPD(PF_38_66, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x02), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, OpSize::i32Bit>},
{OPD(PF_38_66, 0x02), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x03), 1, &OpDispatchBuilder::PHADDS},
{OPD(PF_38_66, 0x03), 1, &OpDispatchBuilder::PHADDS},
{OPD(PF_38_NONE, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_66, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::PHSUB<2>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::PHSUB<2>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::PHSUB<4>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::PHSUB<4>},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_66, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::PSIGN<1>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::PSIGN<1>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::PSIGN<2>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::PSIGN<2>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::PSIGN<4>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::PSIGN<4>},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x10), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, 1>},
{OPD(PF_38_66, 0x14), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, 4>},
{OPD(PF_38_66, 0x15), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, 8>},
{OPD(PF_38_66, 0x10), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, OpSize::i8Bit>},
{OPD(PF_38_66, 0x14), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, OpSize::i32Bit>},
{OPD(PF_38_66, 0x15), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, OpSize::i64Bit>},
{OPD(PF_38_66, 0x17), 1, &OpDispatchBuilder::PTestOp},
{OPD(PF_38_NONE, 0x1C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, 1>},
{OPD(PF_38_66, 0x1C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, 1>},
{OPD(PF_38_NONE, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, 2>},
{OPD(PF_38_66, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, 2>},
{OPD(PF_38_NONE, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, 4>},
{OPD(PF_38_66, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, 4>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<1, 2, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<1, 4, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<1, 8, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<2, 4, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<2, 8, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<4, 8, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::PMULLOp<4, true>},
{OPD(PF_38_66, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, 8>},
{OPD(PF_38_NONE, 0x1C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i8Bit>},
{OPD(PF_38_66, 0x1C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(PF_38_66, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i64Bit>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::PACKUSOp<4>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<1, 2, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<1, 4, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<1, 8, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<2, 4, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<2, 8, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<4, 8, false>},
{OPD(PF_38_66, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, 8>},
{OPD(PF_38_66, 0x38), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, 1>},
{OPD(PF_38_66, 0x39), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, 4>},
{OPD(PF_38_66, 0x3A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, 2>},
{OPD(PF_38_66, 0x3B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, 4>},
{OPD(PF_38_66, 0x3C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, 1>},
{OPD(PF_38_66, 0x3D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, 4>},
{OPD(PF_38_66, 0x3E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, 2>},
{OPD(PF_38_66, 0x3F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, 4>},
{OPD(PF_38_66, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, 4>},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::PACKUSOp<OpSize::i32Bit>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i64Bit>},
{OPD(PF_38_66, 0x38), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i8Bit>},
{OPD(PF_38_66, 0x39), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i32Bit>},
{OPD(PF_38_66, 0x3A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, OpSize::i16Bit>},
{OPD(PF_38_66, 0x3B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, OpSize::i32Bit>},
{OPD(PF_38_66, 0x3C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, OpSize::i8Bit>},
{OPD(PF_38_66, 0x3D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, OpSize::i32Bit>},
{OPD(PF_38_66, 0x3E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, OpSize::i16Bit>},
{OPD(PF_38_66, 0x3F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, OpSize::i32Bit>},
{OPD(PF_38_66, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, OpSize::i32Bit>},
{OPD(PF_38_66, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(PF_38_NONE, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
@@ -6,42 +6,68 @@ namespace FEXCore::IR {
#define OPD(REX, prefix, opcode) ((REX << 9) | (prefix << 8) | opcode)
#define PF_3A_NONE 0
#define PF_3A_66 1
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F3ATable[] = {
{OPD(0, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<4>},
{OPD(0, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<8>},
{OPD(0, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::InsertScalarRound<4>},
{OPD(0, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::InsertScalarRound<8>},
{OPD(0, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::VectorBlend<4>},
{OPD(0, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::VectorBlend<8>},
{OPD(0, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::VectorBlend<2>},
constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatchTableGenH0F3AREX = []<uint16_t REX>() consteval {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> Table[] = {
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::VectorBlend<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::VectorBlend<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::VectorBlend<OpSize::i16Bit>},
{OPD(0, PF_3A_NONE, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(0, PF_3A_66, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(REX, PF_3A_NONE, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(REX, PF_3A_66, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(0, PF_3A_66, 0x14), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, 1>},
{OPD(0, PF_3A_66, 0x15), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, 2>},
{OPD(0, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, 4>},
{OPD(0, PF_3A_66, 0x17), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, 4>},
{OPD(REX, PF_3A_66, 0x14), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x15), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{OPD(REX, PF_3A_66, 0x17), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x20), 1, &OpDispatchBuilder::PINSROp<1>},
{OPD(0, PF_3A_66, 0x21), 1, &OpDispatchBuilder::InsertPSOp},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<4>},
{OPD(0, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<4>},
{OPD(0, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<8>},
{OPD(0, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(REX, PF_3A_66, 0x20), 1, &OpDispatchBuilder::PINSROp<OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x21), 1, &OpDispatchBuilder::InsertPSOp},
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(0, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(0, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(0, PF_3A_66, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(0, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(REX, PF_3A_66, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(0, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
{OPD(REX, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
};
return std::to_array(Table);
};
auto REX0 = OpDispatchTableGenH0F3AREX.template operator()<0>();
auto REX1 = OpDispatchTableGenH0F3AREX.template operator()<1>();
auto concat = []<typename T, size_t N1, size_t N2>(std::array<T, N1> const& lhs,
std::array<T, N2> const& rhs) consteval -> std::array<T, N1 + N2> {
std::array<T, N1 + N2> Table {};
for (size_t i = 0; i < N1; ++i) {
Table[i] = lhs[i];
}
for (size_t i = 0; i < N2; ++i) {
Table[N1 + i] = rhs[i];
}
return Table;
};
return concat(REX0, REX1);
};
constexpr auto OpDispatch_H0F3ATableIgnoreREX = OpDispatchTableGenH0F3A();
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F3ATableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i32Bit>},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(1, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, 8>},
{OPD(1, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<8>},
{OPD(1, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i64Bit>},
{OPD(1, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i64Bit>},
};
#undef PF_3A_NONE
@@ -21,6 +21,11 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_66, 0), 1, &OpDispatchBuilder::SGDTOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_F2, 0), 1, &OpDispatchBuilder::SGDTOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_NONE, 1), 1, &OpDispatchBuilder::SIDTOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_F3, 1), 1, &OpDispatchBuilder::SIDTOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_66, 1), 1, &OpDispatchBuilder::SIDTOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_F2, 1), 1, &OpDispatchBuilder::SIDTOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_NONE, 3), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_F3, 3), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_66, 3), 1, &OpDispatchBuilder::PermissionRestrictedOp},
@@ -36,6 +41,11 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_66, 6), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_F2, 6), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_NONE, 7), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_F3, 7), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_66, 7), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_7, PF_F2, 7), 1, &OpDispatchBuilder::PermissionRestrictedOp},
// GROUP 8
{OPD(FEXCore::X86Tables::TYPE_GROUP_8, PF_NONE, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::BTOp, 1, BTAction::BTNone>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_8, PF_F3, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::BTOp, 1, BTAction::BTNone>},
@@ -66,30 +76,30 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 7), 1, &OpDispatchBuilder::RDPIDOp},
// GROUP 12
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_NONE, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, 2>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_NONE, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, 2>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, 2>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_NONE, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, OpSize::i16Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_NONE, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, OpSize::i16Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, OpSize::i16Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_66, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, 2>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_66, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, 2>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, 2>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_66, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, OpSize::i16Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_66, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, OpSize::i16Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_12, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, OpSize::i16Bit>},
// GROUP 13
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_NONE, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, 4>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_NONE, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, 4>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, 4>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_NONE, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, OpSize::i32Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_NONE, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, OpSize::i32Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, OpSize::i32Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_66, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, 4>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_66, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, 4>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, 4>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_66, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, OpSize::i32Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_66, 4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAIOp, OpSize::i32Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_13, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, OpSize::i32Bit>},
// GROUP 14
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_NONE, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, 8>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, 8>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_NONE, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, OpSize::i64Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, OpSize::i64Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_66, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, 8>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_66, 2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLI, OpSize::i64Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_66, 3), 1, &OpDispatchBuilder::PSRLDQ},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, 8>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLLI, OpSize::i64Bit>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_14, PF_66, 7), 1, &OpDispatchBuilder::PSLLDQ},
// GROUP 15
@@ -5,6 +5,7 @@
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable[] = {
// Instructions
{0x03, 1, &OpDispatchBuilder::LSLOp},
{0x06, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x07, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x0B, 1, &OpDispatchBuilder::INTOp},
@@ -44,104 +45,104 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xBE, 2, &OpDispatchBuilder::MOVSXOp},
{0xC0, 2, &OpDispatchBuilder::XADDOp},
{0xC3, 1, &OpDispatchBuilder::MOVGPRNTOp},
{0xC4, 1, &OpDispatchBuilder::PINSROp<2>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, 2>},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC8, 8, &OpDispatchBuilder::BSWAPOp},
// SSE
{0x10, 2, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x12, 2, &OpDispatchBuilder::MOVLPOp},
{0x14, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 4>},
{0x15, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 4>},
{0x14, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x15, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::InsertMMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<4, false, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<4, false, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<4>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, 4>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, 4>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRSQRT, 4>},
{0x53, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECP, 4>},
{0x54, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, 16>},
{0x55, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, 8>},
{0x56, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, 16>},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i32Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRSQRT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECP, OpSize::i32Bit>},
{0x54, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
{0x55, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, OpSize::i64Bit>},
{0x56, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, OpSize::i128Bit>},
{0x57, 1, &OpDispatchBuilder::VectorXOROp},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, 4>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, 4>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, 8, 4, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<4, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, 4>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, 4>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, 4>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, 4>},
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 1>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 2>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 4>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<2>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, 1>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, 2>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, 4>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<2>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 1>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 2>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 4>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<4>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x70, 1, &OpDispatchBuilder::PSHUFW8ByteOp},
{0x74, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, 1>},
{0x75, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, 2>},
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, 4>},
{0x74, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i8Bit>},
{0x75, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i16Bit>},
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i32Bit>},
{0x77, 1, &OpDispatchBuilder::X87EMMS},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<4>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, 4>},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i32Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i32Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, 2>},
{0xD2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, 4>},
{0xD3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, 8>},
{0xD4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 8>},
{0xD5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, 2>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
{0xD2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i32Bit>},
{0xD3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i64Bit>},
{0xD4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i64Bit>},
{0xD5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, OpSize::i16Bit>},
{0xD7, 1, &OpDispatchBuilder::MOVMSKOpOne}, // PMOVMSKB
{0xD8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, 1>},
{0xD9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, 2>},
{0xDA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, 1>},
{0xDB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, 8>},
{0xDC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, 1>},
{0xDD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, 2>},
{0xDE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, 1>},
{0xDF, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, 8>},
{0xE0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, 1>},
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, 2>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, 4>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, 2>},
{0xD8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, OpSize::i8Bit>},
{0xD9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, OpSize::i16Bit>},
{0xDA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, OpSize::i8Bit>},
{0xDB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, OpSize::i64Bit>},
{0xDC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, OpSize::i8Bit>},
{0xDD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, OpSize::i16Bit>},
{0xDE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, OpSize::i8Bit>},
{0xDF, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, OpSize::i64Bit>},
{0xE0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i8Bit>},
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, 1>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, 2>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, 2>},
{0xEB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, 8>},
{0xEC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, 1>},
{0xED, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, 2>},
{0xEE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, 2>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
{0xEB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, OpSize::i64Bit>},
{0xEC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, OpSize::i8Bit>},
{0xED, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, OpSize::i16Bit>},
{0xEE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, OpSize::i16Bit>},
{0xEF, 1, &OpDispatchBuilder::VectorXOROp},
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, 2>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, 4>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, 8>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<4, false>},
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
{0xF8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 1>},
{0xF9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 2>},
{0xFA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 4>},
{0xFB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 8>},
{0xFC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 1>},
{0xFD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 2>},
{0xFE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 4>},
{0xF8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i8Bit>},
{0xF9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i16Bit>},
{0xFA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i32Bit>},
{0xFB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i64Bit>},
{0xFC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i8Bit>},
{0xFD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i16Bit>},
{0xFE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
// FEX reserved instructions
{0x37, 1, &OpDispatchBuilder::CallbackReturnOp},
@@ -151,21 +152,21 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0x10, 2, &OpDispatchBuilder::MOVSSOp},
{0x12, 1, &OpDispatchBuilder::VMOVSLDUPOp},
{0x16, 1, &OpDispatchBuilder::VMOVSHDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<4>},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<4, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<4, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, 4>},
{0x52, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, 4>},
{0x53, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, 4>},
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, 4>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, 4>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<8, 4>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, false>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, 4>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, 4>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, 4>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, 4>},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i64Bit, OpSize::i32Bit>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{0x6F, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, false>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQOp, OpDispatchBuilder::VectorOpType::SSE>},
@@ -173,142 +174,142 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xB8, 1, &OpDispatchBuilder::PopcountOp},
{0xBC, 1, &OpDispatchBuilder::TZCNT},
{0xBD, 1, &OpDispatchBuilder::LZCNT},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<4>},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<4, true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryRepNEModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSDOp},
{0x12, 1, &OpDispatchBuilder::MOVDDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<8>},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<8, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<8, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, 8>},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
// x52 = Invalid
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, 8>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, 8>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<4, 8>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, 8>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, 8>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, 8>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, 8>},
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i32Bit, OpSize::i64Bit>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, true>},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, 4>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<4>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<4>},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<false>},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<8>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<8, true, true>},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i64Bit>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0xF0, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryOpSizeModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x12, 2, &OpDispatchBuilder::MOVLPOp},
{0x14, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 8>},
{0x15, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 8>},
{0x14, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i64Bit>},
{0x15, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i64Bit>},
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<8, true, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<8, true, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<8>},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, 8>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, 8>},
{0x54, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, 16>},
{0x55, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, 8>},
{0x56, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, 16>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i64Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i64Bit>},
{0x54, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
{0x55, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, OpSize::i64Bit>},
{0x56, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, OpSize::i128Bit>},
{0x57, 1, &OpDispatchBuilder::VectorXOROp},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, 8>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, 8>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, 4, 8, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, true>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, 8>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, 8>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, 8>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, 8>},
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 1>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 2>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 4>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<2>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, 1>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, 2>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, 4>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<2>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 1>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 2>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 4>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<4>},
{0x6C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, 8>},
{0x6D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, 8>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, OpSize::i64Bit>},
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x6C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i64Bit>},
{0x6D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i64Bit>},
{0x6E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
{0x6F, 1, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x70, 1, &OpDispatchBuilder::PSHUFDOp},
{0x74, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, 1>},
{0x75, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, 2>},
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, 4>},
{0x74, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i8Bit>},
{0x75, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i16Bit>},
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i32Bit>},
{0x78, 1, nullptr}, // GROUP 17
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, 8>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<8>},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i64Bit>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
{0x7F, 1, &OpDispatchBuilder::MOVVectorAlignedOp},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<8>},
{0xC4, 1, &OpDispatchBuilder::PINSROp<2>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, 2>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, 8>},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i64Bit>},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i64Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<8>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, 2>},
{0xD2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, 4>},
{0xD3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, 8>},
{0xD4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 8>},
{0xD5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, 2>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i64Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
{0xD2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i32Bit>},
{0xD3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i64Bit>},
{0xD4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i64Bit>},
{0xD5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, OpSize::i16Bit>},
{0xD6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQOp, OpDispatchBuilder::VectorOpType::SSE>},
{0xD7, 1, &OpDispatchBuilder::MOVMSKOpOne}, // PMOVMSKB
{0xD8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, 1>},
{0xD9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, 2>},
{0xDA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, 1>},
{0xDB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, 16>},
{0xDC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, 1>},
{0xDD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, 2>},
{0xDE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, 1>},
{0xDF, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, 8>},
{0xE0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, 1>},
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, 2>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, 4>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, 2>},
{0xD8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, OpSize::i8Bit>},
{0xD9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQSUB, OpSize::i16Bit>},
{0xDA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMIN, OpSize::i8Bit>},
{0xDB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
{0xDC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, OpSize::i8Bit>},
{0xDD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUQADD, OpSize::i16Bit>},
{0xDE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VUMAX, OpSize::i8Bit>},
{0xDF, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VANDN, OpSize::i64Bit>},
{0xE0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i8Bit>},
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<8, true, false>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, 1>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, 2>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, 2>},
{0xEB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, 16>},
{0xEC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, 1>},
{0xED, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, 2>},
{0xEE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, 2>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
{0xEB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VOR, OpSize::i128Bit>},
{0xEC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, OpSize::i8Bit>},
{0xED, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQADD, OpSize::i16Bit>},
{0xEE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMAX, OpSize::i16Bit>},
{0xEF, 1, &OpDispatchBuilder::VectorXOROp},
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, 2>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, 4>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, 8>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<4, false>},
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
{0xF8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 1>},
{0xF9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 2>},
{0xFA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 4>},
{0xFB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, 8>},
{0xFC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 1>},
{0xFD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 2>},
{0xFE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, 4>},
{0xF8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i8Bit>},
{0xF9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i16Bit>},
{0xFA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i32Bit>},
{0xFB, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSUB, OpSize::i64Bit>},
{0xFC, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i8Bit>},
{0xFD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i16Bit>},
{0xFE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable_64[] = {
File diff suppressed because it is too large. Load diff
@@ -16,6 +16,7 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/FPState.h>
#include <cmath>
#include <stddef.h>
#include <stdint.h>
@@ -26,7 +27,7 @@ class OrderedNode;
Ref OpDispatchBuilder::GetX87Top() {
// Yes, we are storing 3 bits in a single flag register.
// Deal with it
return _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
return _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
Ref OpDispatchBuilder::GetX87Tag(Ref Value, Ref AbridgedFTW) {
@@ -39,7 +40,7 @@ Ref OpDispatchBuilder::GetX87Tag(Ref Value, Ref AbridgedFTW) {
void OpDispatchBuilder::SetX87FTW(Ref FTW) {
Ref X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
Ref NewAbridgedFTW;
Ref NewAbridgedFTW {};
for (int i = 0; i < 8; i++) {
Ref RegTag = _Bfe(OpSize::i32Bit, 2, i * 2, FTW);
@@ -56,17 +57,17 @@ void OpDispatchBuilder::SetX87FTW(Ref FTW) {
}
void OpDispatchBuilder::SetX87Top(Ref Value) {
_StoreContext(1, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
// Float LoaD operation with memory operand
void OpDispatchBuilder::FLD(OpcodeArgs, size_t Width) {
size_t ReadWidth = (Width == 80) ? 16 : Width / 8;
void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], Width, Op->Flags);
Ref ConvertedData = Data;
// Convert to 80bit float
if (Width == 32 || Width == 64) {
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
ConvertedData = _F80CVTTo(Data, ReadWidth);
}
_PushStack(ConvertedData, Data, ReadWidth, true);
@@ -79,31 +80,31 @@ void OpDispatchBuilder::FLDFromStack(OpcodeArgs) {
void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Data, 16, true);
_PushStack(ConvertedData, Data, OpSize::i128Bit, true);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
Ref converted = _F80BCDStore(_ReadStackValue(0));
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, 10, 1);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
_PopStackDestroy();
}
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant Constant) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(16, Constant);
_PushStack(Data, Data, 16, true);
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, Constant);
_PushStack(Data, Data, OpSize::i128Bit, true);
}
void OpDispatchBuilder::FILD(OpcodeArgs) {
size_t ReadWidth = GetSrcSize(Op);
const auto ReadWidth = OpSizeFromSrc(Op);
// Read from memory
Ref Data = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
// Sign extend to 64bits
if (ReadWidth != 8) {
Data = _Sbfe(OpSize::i64Bit, ReadWidth * 8, 0, Data);
if (ReadWidth != OpSize::i64Bit) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
// We're about to clobber flags to grab the sign, so save NZCV.
@@ -123,14 +124,29 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
auto zeroed_exponent = _Select(COND_EQ, absolute, zero, zero, adjusted_exponent);
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VCastFromGPR(16, 8, shifted);
ConvertedData = _VInsElement(16, 8, 1, 0, ConvertedData, _VCastFromGPR(16, 8, upper));
Ref ConvertedData = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, shifted);
ConvertedData = _VInsElement(OpSize::i128Bit, OpSize::i64Bit, 1, 0, ConvertedData, _VCastFromGPR(OpSize::i128Bit, OpSize::i64Bit, upper));
_PushStack(ConvertedData, Data, ReadWidth, false);
}
void OpDispatchBuilder::FST(OpcodeArgs, size_t Width) {
Ref Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
_StoreStackMemory(Mem, OpSize::i128Bit, true, Width / 8);
void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
// Ref Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
// FIXME: Is TSO relevant for x87?
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
// Index scale is a power of 2?
LOGMAN_THROW_A_FMT(A.IndexScale > 0 && (A.IndexScale & (A.IndexScale - 1)) == 0, "Invalid index scale");
Ref Addr = A.Base ? A.Base : _Constant(0);
if (A.Index) {
Ref ScaledIndex = A.Index;
if (A.IndexScale > 1) {
ScaledIndex = _Lshl(A.AddrSize, ScaledIndex, _Constant(std::log2(A.IndexScale)));
}
Addr = _Add(A.AddrSize, Addr, ScaledIndex);
}
_StoreStackMem(OpSize::i128Bit, Width, Addr, _Constant(A.Offset), /*Float=*/true);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
_PopStackDestroy();
}
@@ -149,18 +165,18 @@ void OpDispatchBuilder::FSTToStack(OpcodeArgs) {
// Store integer to memory (possibly with truncation)
void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
auto Size = GetSrcSize(Op);
const auto Size = OpSizeFromSrc(Op);
Ref Data = _ReadStackValue(0);
Data = _F80CVTInt(Size, Data, Truncate);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Data, Size, 1);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
}
}
void OpDispatchBuilder::FADD(OpcodeArgs, size_t Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FADD(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) { // Implicit argument case
auto Offset = Op->OP & 7;
auto St0 = 0;
@@ -175,22 +191,22 @@ void OpDispatchBuilder::FADD(OpcodeArgs, size_t Width, bool Integer, OpDispatchB
return;
}
LOGMAN_THROW_A_FMT(Width != 80, "No 80-bit floats from memory");
LOGMAN_THROW_A_FMT(Width != OpSize::f80Bit, "No 80-bit floats from memory");
// We have one memory argument
Ref Arg {};
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = _F80CVTToInt(Arg, Width / 8);
Arg = _F80CVTToInt(Arg, Width);
} else {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = _F80CVTTo(Arg, Width / 8);
Arg = _F80CVTTo(Arg, Width);
}
// top of stack is at offset zero
_F80AddValue(0, Arg);
}
void OpDispatchBuilder::FMUL(OpcodeArgs, size_t Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FMUL(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) { // Implicit argument case
auto offset = Op->OP & 7;
auto st0 = 0;
@@ -205,15 +221,15 @@ void OpDispatchBuilder::FMUL(OpcodeArgs, size_t Width, bool Integer, OpDispatchB
return;
}
LOGMAN_THROW_A_FMT(Width != 80, "No 80-bit floats from memory");
LOGMAN_THROW_A_FMT(Width != OpSize::f80Bit, "No 80-bit floats from memory");
// We have one memory argument
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = _F80CVTToInt(arg, Width / 8);
arg = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = _F80CVTTo(arg, Width / 8);
arg = _F80CVTTo(arg, Width);
}
// top of stack is at offset zero
@@ -224,11 +240,11 @@ void OpDispatchBuilder::FMUL(OpcodeArgs, size_t Width, bool Integer, OpDispatchB
}
}
void OpDispatchBuilder::FDIV(OpcodeArgs, size_t Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FDIV(OpcodeArgs, IR::OpSize Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) {
const auto Offset = Op->OP & 7;
const auto St0 = 0;
const auto Result = (ResInST0 == OpResult::RES_STI) ? Offset : St0;
const uint8_t Offset = Op->OP & 7;
const uint8_t St0 = 0;
const uint8_t Result = (ResInST0 == OpResult::RES_STI) ? Offset : St0;
if (Reverse ^ (ResInST0 == OpResult::RES_STI)) {
_F80DivStack(Result, Offset, St0);
@@ -242,15 +258,15 @@ void OpDispatchBuilder::FDIV(OpcodeArgs, size_t Width, bool Integer, bool Revers
return;
}
LOGMAN_THROW_A_FMT(Width != 80, "No 80-bit floats from memory");
LOGMAN_THROW_A_FMT(Width != OpSize::f80Bit, "No 80-bit floats from memory");
// We have one memory argument
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = _F80CVTToInt(arg, Width / 8);
arg = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = _F80CVTTo(arg, Width / 8);
arg = _F80CVTTo(arg, Width);
}
// top of stack is at offset zero
@@ -265,7 +281,7 @@ void OpDispatchBuilder::FDIV(OpcodeArgs, size_t Width, bool Integer, bool Revers
}
}
void OpDispatchBuilder::FSUB(OpcodeArgs, size_t Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FSUB(OpcodeArgs, IR::OpSize Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) {
const auto Offset = Op->OP & 7;
const auto St0 = 0;
@@ -283,15 +299,15 @@ void OpDispatchBuilder::FSUB(OpcodeArgs, size_t Width, bool Integer, bool Revers
return;
}
LOGMAN_THROW_A_FMT(Width != 80, "No 80-bit floats from memory");
LOGMAN_THROW_A_FMT(Width != OpSize::f80Bit, "No 80-bit floats from memory");
// We have one memory argument
Ref Arg {};
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = _F80CVTToInt(Arg, Width / 8);
Arg = _F80CVTToInt(Arg, Width);
} else {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = _F80CVTTo(Arg, Width / 8);
Arg = _F80CVTTo(Arg, Width);
}
// top of stack is at offset zero
@@ -342,42 +358,42 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
// Before we store anything we need to sync our stack to the registers.
_SyncStackToSlow();
auto Size = GetDstSize(Op);
const auto Size = OpSizeFromSrc(Op);
Ref Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
{
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
auto ZeroConst = _Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
}
}
@@ -400,26 +416,27 @@ Ref OpDispatchBuilder::ReconstructX87StateFromFSW_Helper(Ref FSW) {
void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
_StackForceSlow();
auto Size = GetSrcSize(Op);
const auto Size = OpSizeFromSrc(Op);
Ref Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(IR::OpSizeToSize(Size) * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(IR::OpSizeToSize(Size) * 2));
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
}
}
void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
_SyncStackToSlow();
// 14 bytes for 16bit
// 2 Bytes : FCW
// 2 Bytes : FSW
@@ -438,60 +455,66 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
// 2 bytes : Opcode
// 4 bytes : data pointer offset
// 4 bytes : data pointer selector
const auto Size = GetDstSize(Op);
const auto Size = OpSizeFromDst(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
Ref Top = GetX87Top();
{
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
auto ZeroConst = _Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
}
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
const auto LoadSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
auto data = _LoadContextIndexed(Top, 16, MMBaseOffset(), 16, FPRClass);
_StoreMem(FPRClass, 16, data, Mem, _Constant((Size * 7) + (10 * i)), 1, MEM_OFFSET_SXTX, 1);
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
_StoreMem(FPRClass, OpSize::i128Bit, data, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
}
// The final st(7) needs a bit of special handling here
auto data = _LoadContextIndexed(Top, 16, MMBaseOffset(), 16, FPRClass);
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
_StoreMem(FPRClass, 8, data, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
auto topBytes = _VDupElement(16, 2, data, 4);
_StoreMem(FPRClass, 2, topBytes, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
_StoreMem(FPRClass, OpSize::i64Bit, data, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
auto topBytes = _VDupElement(OpSize::i128Bit, OpSize::i16Bit, data, 4);
_StoreMem(FPRClass, OpSize::i16Bit, topBytes, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
// reset to default
FNINIT(Op);
@@ -499,17 +522,27 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
_StackForceSlow();
const auto Size = GetSrcSize(Op);
const auto Size = OpSizeFromSrc(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
if (ReducedPrecisionMode) {
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = NewFCW;
auto roundShift = _Constant(10);
auto roundMask = _Constant(3);
roundingMode = _Lshr(OpSize::i32Bit, roundingMode, roundShift);
roundingMode = _And(OpSize::i32Bit, roundingMode, roundMask);
_SetRoundingMode(roundingMode, false, roundingMode);
}
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1);
Ref Top = ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
}
auto OneConst = _Constant(1);
@@ -517,15 +550,18 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
auto low = _Constant(~0ULL);
auto high = _Constant(0xFFFF);
Ref Mask = _VCastFromGPR(16, 8, low);
Mask = _VInsGPR(16, 8, 1, Mask, high);
Ref Mask = _VCastFromGPR(OpSize::i128Bit, OpSize::i64Bit, low);
Mask = _VInsGPR(OpSize::i128Bit, OpSize::i64Bit, 1, Mask, high);
const auto StoreSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMem(FPRClass, 16, Mem, _Constant((Size * 7) + (10 * i)), 1, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMem(FPRClass, OpSize::i128Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
// Mask off the top bits
Reg = _VAnd(16, 16, Reg, Mask);
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
Reg = _VAnd(OpSize::i128Bit, OpSize::i128Bit, Reg, Mask);
if (ReducedPrecisionMode) {
// Convert to double precision
Reg = _F80CVT(OpSize::i64Bit, Reg);
}
_StoreContextIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
}
@@ -534,29 +570,31 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
Ref Reg = _LoadMem(FPRClass, 8, Mem, _Constant((Size * 7) + (10 * 7)), 1, MEM_OFFSET_SXTX, 1);
Ref RegHigh = _LoadMem(FPRClass, 2, Mem, _Constant((Size * 7) + (10 * 7) + 8), 1, MEM_OFFSET_SXTX, 1);
Reg = _VInsElement(16, 2, 4, 0, Reg, RegHigh);
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
Ref Reg = _LoadMem(FPRClass, OpSize::i64Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref RegHigh =
_LoadMem(FPRClass, OpSize::i16Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Reg = _VInsElement(OpSize::i128Bit, OpSize::i16Bit, 4, 0, Reg, RegHigh);
if (ReducedPrecisionMode) {
Reg = _F80CVT(OpSize::i64Bit, Reg); // Convert to double precision
}
_StoreContextIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
}
// Load / Store Control Word
void OpDispatchBuilder::X87FSTCW(OpcodeArgs) {
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
StoreResult(GPRClass, Op, FCW, -1);
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
StoreResult(GPRClass, Op, FCW, OpSize::iInvalid);
}
void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
// FIXME: Because loading control flags will affect several instructions in fast path, we might have
// to switch for now to slow mode whenever these are manually changed.
// Remove the next line and try DF_04.asm in fast path.
_StackForceSlow();
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
void OpDispatchBuilder::FXCH(OpcodeArgs) {
uint8_t Offset = Op->OP & 7;
// fxch st0, st0 is for us essentially a nop
@@ -569,15 +607,15 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
void OpDispatchBuilder::X87FYL2X(OpcodeArgs, bool IsFYL2XP1) {
if (IsFYL2XP1) {
// create an add between top of stack and 1.
Ref One = ReducedPrecisionMode ? _VCastFromGPR(8, 8, _Constant(0x3FF0000000000000)) :
LoadAndCacheNamedVectorConstant(16, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(0x3FF0000000000000)) :
LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
_F80AddValue(0, One);
}
_F80FYL2XStack();
}
void OpDispatchBuilder::FCOMI(OpcodeArgs, size_t Width, bool Integer, OpDispatchBuilder::FCOMIFlags WhichFlags, bool PopTwice) {
void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::FCOMIFlags WhichFlags, bool PopTwice) {
Ref arg {};
Ref b {};
@@ -587,15 +625,17 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, size_t Width, bool Integer, OpDispatch
uint8_t Offset = Op->OP & 7;
Res = _F80CmpStack(Offset);
} else {
// Memory arg
if (Width == 16 || Width == 32 || Width == 64) {
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
// Memory arg
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
b = _F80CVTToInt(arg, Width / 8);
b = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
b = _F80CVTTo(arg, Width / 8);
b = _F80CVTTo(arg, Width);
}
} else {
FEX_UNREACHABLE;
}
Res = _F80CmpValue(b);
}
@@ -612,10 +652,7 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, size_t Width, bool Integer, OpDispatch
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
} else {
// Invalidate deferred flags early
// OF, SF, AF, PF all undefined
InvalidateDeferredFlags();
SetCFDirect(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_RAW_LOC>(HostFlag_ZF);
@@ -675,7 +712,6 @@ void OpDispatchBuilder::X87ModifySTP(OpcodeArgs, bool Inc) {
// Optionally we can pass a pre calculated value for Top, otherwise we calculate it
// during the function runtime.
Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
// Start with the top value
auto Top = T ? T : GetX87Top();
Ref FSW = _Lshl(OpSize::i64Bit, Top, _Constant(11));
@@ -700,18 +736,21 @@ Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
// There's no load Status Word instruction but you can load it through frstor
// or fldenv.
void OpDispatchBuilder::X87FNSTSW(OpcodeArgs) {
Ref TopValue = _SyncStackToSlow();
Ref StatusWord = ReconstructFSW_Helper(TopValue);
StoreResult(GPRClass, Op, StatusWord, -1);
StoreResult(GPRClass, Op, StatusWord, OpSize::iInvalid);
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
auto Zero = _Constant(0);
if (ReducedPrecisionMode) {
_SetRoundingMode(Zero, false, Zero);
}
// Init FCW to 0x037F
auto NewFCW = _Constant(16, 0x037F);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFCW = _Constant(OpSize::i16Bit, 0x037F);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
// Set top to zero
SetX87Top(Zero);
@@ -728,13 +767,11 @@ void OpDispatchBuilder::FNINIT(OpcodeArgs) {
}
void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
_InvalidateStack(Op->OP & 7);
}
void OpDispatchBuilder::X87EMMS(OpcodeArgs) {
// Tags all get set to 0b11
_InvalidateStack(0xff);
}
@@ -776,13 +813,14 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
auto AllOneConst = _Constant(0xffff'ffff'ffff'ffffull);
Ref SrcCond = SelectCC(CC, OpSize::i64Bit, AllOneConst, ZeroConst);
Ref VecCond = _VDupFromGPR(16, 8, SrcCond);
_F80VBSLStack(16, VecCond, Op->OP & 7, 0);
Ref VecCond = _VDupFromGPR(OpSize::i128Bit, OpSize::i64Bit, SrcCond);
_F80VBSLStack(OpSize::i128Bit, VecCond, Op->OP & 7, 0);
}
void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
auto a = _ReadStackValue(0);
Ref Result = ReducedPrecisionMode ? _VExtractToGPR(8, 8, a, 0) : _VExtractToGPR(16, 8, a, 1);
Ref Result =
ReducedPrecisionMode ? _VExtractToGPR(OpSize::i64Bit, OpSize::i64Bit, a, 0) : _VExtractToGPR(OpSize::i128Bit, OpSize::i64Bit, a, 1);
// Extract the sign bit
Result = ReducedPrecisionMode ? _Bfe(OpSize::i64Bit, 1, 63, Result) : _Bfe(OpSize::i64Bit, 1, 15, Result);
@@ -804,4 +842,14 @@ void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(C3);
}
void OpDispatchBuilder::X87FXTRACT(OpcodeArgs) {
auto Top = _ReadStackValue(0);
_PopStackDestroy();
auto Exp = _F80XTRACT_EXP(Top);
auto Sig = _F80XTRACT_SIG(Top);
_PushStack(Exp, Exp, OpSize::f80Bit, true);
_PushStack(Sig, Sig, OpSize::f80Bit, true);
}
} // namespace FEXCore::IR
@@ -8,6 +8,7 @@ $end_info$
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/IR/IR.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
@@ -22,38 +23,28 @@ class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::FNINITF64(OpcodeArgs) {
// Init host rounding mode to zero
auto Zero = _Constant(0);
_SetRoundingMode(Zero, false, Zero);
// Call generic version
FNINIT(Op);
}
void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
_StackForceSlow();
const auto Size = GetSrcSize(Op);
const auto Size = OpSizeFromSrc(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size)), Size, MEM_OFFSET_SXTX, 1);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
}
}
void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
_StackForceSlow();
@@ -62,82 +53,94 @@ void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
// F64 ops
// Float load op with memory operand
void OpDispatchBuilder::FLDF64(OpcodeArgs, size_t Width) {
size_t ReadWidth = (Width == 80) ? 16 : Width / 8;
void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
// Convert to 64bit float
Ref ConvertedData = Data;
if (Width == 32) {
ConvertedData = _Float_FToF(8, 4, Data);
} else if (Width == 80) {
ConvertedData = _F80CVT(8, Data);
if (Width == OpSize::i32Bit) {
ConvertedData = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, Data);
} else if (Width == OpSize::f80Bit) {
ConvertedData = _F80CVT(OpSize::i64Bit, Data);
}
_PushStack(ConvertedData, Data, ReadWidth, true);
}
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
// Read from memory
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::i128Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(8, ConvertedData);
_PushStack(ConvertedData, Data, 8, true);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Data, OpSize::i64Bit, true);
}
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
Ref converted = _F80CVTTo(_ReadStackValue(0), 8);
Ref converted = _F80CVTTo(_ReadStackValue(0), OpSize::i64Bit);
converted = _F80BCDStore(converted);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, 10, 1);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
_PopStackDestroy();
}
void OpDispatchBuilder::FLDF64_Const(OpcodeArgs, uint64_t Num) {
auto Data = _VCastFromGPR(8, 8, _Constant(Num));
_PushStack(Data, Data, 8, true);
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(Num));
_PushStack(Data, Data, OpSize::i64Bit, true);
}
void OpDispatchBuilder::FILDF64(OpcodeArgs) {
size_t ReadWidth = GetSrcSize(Op);
const auto ReadWidth = OpSizeFromSrc(Op);
// Read from memory
Ref Data = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
if (ReadWidth == 2) {
Data = _Sbfe(OpSize::i64Bit, ReadWidth * 8, 0, Data);
if (ReadWidth == OpSize::i16Bit) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
auto ConvertedData = _Float_FromGPR_S(8, ReadWidth == 4 ? 4 : 8, Data);
auto ConvertedData = _Float_FromGPR_S(OpSize::i64Bit, ReadWidth == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, Data);
_PushStack(ConvertedData, Data, ReadWidth, false);
}
void OpDispatchBuilder::FSTF64(OpcodeArgs, size_t Width) {
Ref Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
_StoreStackMemory(Mem, OpSize::i64Bit, true, Width / 8);
void OpDispatchBuilder::FSTF64(OpcodeArgs, IR::OpSize Width) {
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
// Index scale is a power of 2?
LOGMAN_THROW_A_FMT(A.IndexScale > 0 && (A.IndexScale & (A.IndexScale - 1)) == 0, "Invalid index scale");
Ref Addr = A.Base ? A.Base : _Constant(0);
if (A.Index) {
Ref ScaledIndex = A.Index;
if (A.IndexScale > 1) {
ScaledIndex = _Lshl(A.AddrSize, ScaledIndex, _Constant(std::log2(A.IndexScale)));
}
Addr = _Add(A.AddrSize, Addr, ScaledIndex);
}
_StoreStackMem(OpSize::i64Bit, Width, Addr, _Constant(A.Offset), /*Float=*/true);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
_PopStackDestroy();
}
}
void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
auto Size = GetSrcSize(Op);
const auto Size = OpSizeFromSrc(Op);
Ref data = _ReadStackValue(0);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == 4 ? 4 : 8, 8, data);
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
} else {
data = _Float_ToGPR_S(Size == 4 ? 4 : 8, 8, data);
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, 1);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
}
}
void OpDispatchBuilder::FADDF64(OpcodeArgs, size_t Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FADDF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) { // Implicit argument case
auto Offset = Op->OP & 7;
auto St0 = 0;
@@ -157,15 +160,17 @@ void OpDispatchBuilder::FADDF64(OpcodeArgs, size_t Width, bool Integer, OpDispat
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if (Width == 16) {
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(8, Width == 64 ? 8 : 4, arg);
} else if (Width == 32) {
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(8, 4, arg);
} else if (Width == 64) {
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
} else {
FEX_UNREACHABLE;
}
// top of stack is at offset zero
@@ -173,7 +178,7 @@ void OpDispatchBuilder::FADDF64(OpcodeArgs, size_t Width, bool Integer, OpDispat
}
// FIXME: following is very similar to FADDF64
void OpDispatchBuilder::FMULF64(OpcodeArgs, size_t Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FMULF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) { // Implicit argument case
auto offset = Op->OP & 7;
auto st0 = 0;
@@ -193,15 +198,17 @@ void OpDispatchBuilder::FMULF64(OpcodeArgs, size_t Width, bool Integer, OpDispat
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if (Width == 16) {
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(8, Width == 64 ? 8 : 4, arg);
} else if (Width == 32) {
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(8, 4, arg);
} else if (Width == 64) {
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
} else {
FEX_UNREACHABLE;
}
// top of stack is at offset zero
@@ -212,7 +219,7 @@ void OpDispatchBuilder::FMULF64(OpcodeArgs, size_t Width, bool Integer, OpDispat
}
}
void OpDispatchBuilder::FDIVF64(OpcodeArgs, size_t Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FDIVF64(OpcodeArgs, IR::OpSize Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) {
const auto offset = Op->OP & 7;
const auto st0 = 0;
@@ -240,19 +247,21 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs, size_t Width, bool Integer, bool Rev
// We have one memory argument
Ref Arg {};
if (Width == 16 || Width == 32 || Width == 64) {
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if (Width == 16) {
if (Width == OpSize::i16Bit) {
Arg = _Sbfe(OpSize::i64Bit, 16, 0, Arg);
}
Arg = _Float_FromGPR_S(8, Width == 64 ? 8 : 4, Arg);
} else if (Width == 32) {
Arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, Arg);
} else if (Width == OpSize::i32Bit) {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = _Float_FToF(8, 4, Arg);
} else if (Width == 64) {
Arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, Arg);
} else if (Width == OpSize::i64Bit) {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
}
// top of stack is at offset zero
@@ -267,7 +276,7 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs, size_t Width, bool Integer, bool Rev
}
}
void OpDispatchBuilder::FSUBF64(OpcodeArgs, size_t Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
void OpDispatchBuilder::FSUBF64(OpcodeArgs, IR::OpSize Width, bool Integer, bool Reverse, OpDispatchBuilder::OpResult ResInST0) {
if (Op->Src[0].IsNone()) {
const auto Offset = Op->OP & 7;
const auto St0 = 0;
@@ -295,19 +304,21 @@ void OpDispatchBuilder::FSUBF64(OpcodeArgs, size_t Width, bool Integer, bool Rev
// We have one memory argument
Ref arg {};
if (Width == 16 || Width == 32 || Width == 64) {
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if (Width == 16) {
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(8, Width == 64 ? 8 : 4, arg);
} else if (Width == 32) {
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(8, 4, arg);
} else if (Width == 64) {
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
}
// top of stack is at offset zero
@@ -328,11 +339,10 @@ void OpDispatchBuilder::FTSTF64(OpcodeArgs) {
// Now we do our comparison.
_F80StackTest(0);
PossiblySetNZCVBits = ~0;
ConvertNZCVToX87();
}
void OpDispatchBuilder::FCOMIF64(OpcodeArgs, size_t Width, bool Integer, OpDispatchBuilder::FCOMIFlags WhichFlags, bool PopTwice) {
void OpDispatchBuilder::FCOMIF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::FCOMIFlags WhichFlags, bool PopTwice) {
Ref arg {};
Ref b {};
@@ -340,22 +350,22 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs, size_t Width, bool Integer, OpDispa
// Implicit arg
uint8_t offset = Op->OP & 7;
b = _ReadStackValue(offset);
} else {
} else if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
// Memory arg
if (Width == 16 || Width == 32 || Width == 64) {
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if (Width == 16) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(8, Width == 64 ? 8 : 4, arg);
} else if (Width == 32) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
b = _Float_FToF(8, 4, arg);
} else if (Width == 64) {
b = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
b = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
b = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
}
if (WhichFlags == FCOMIFlags::FLAGS_X87) {
@@ -363,7 +373,6 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs, size_t Width, bool Integer, OpDispa
GetNZCV();
_F80CmpValue(b);
PossiblySetNZCVBits = ~0;
ConvertNZCVToX87();
} else {
HandleNZCVWrite();
@@ -379,144 +388,37 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs, size_t Width, bool Integer, OpDispa
}
}
// This function converts to F80 on save for compatibility
void OpDispatchBuilder::X87FNSAVEF64(OpcodeArgs) {
_SyncStackToSlow();
// 14 bytes for 16bit
// 2 Bytes : FCW
// 2 Bytes : FSW
// 2 bytes : FTW
// 2 bytes : Instruction offset
// 2 bytes : Instruction CS selector
// 2 bytes : Data offset
// 2 bytes : Data selector
void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
// Split node into SIG and EXP while handling the special zero case.
// i.e. if val == 0.0, then sig = 0.0, exp = -inf
// if val == -0.0, then sig = -0.0, exp = -inf
// otherwise we just extract the 64-bit sig and exp as normal.
Ref Node = _ReadStackValue(0);
// 28 bytes for 32bit
// 4 bytes : FCW
// 4 bytes : FSW
// 4 bytes : FTW
// 4 bytes : Instruction pointer
// 2 bytes : instruction pointer selector
// 2 bytes : Opcode
// 4 bytes : data pointer offset
// 4 bytes : data pointer selector
Ref Gpr = _VExtractToGPR(OpSize::i64Bit, OpSize::i64Bit, Node, 0);
const auto Size = GetDstSize(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
Ref Top = GetX87Top();
{
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
// zero case
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1); }
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = _Sub(OpSize::i64Bit, ExpNZ, _Constant(1023));
Ref ExpNZV = _Float_FromGPR_S(OpSize::i64Bit, OpSize::i64Bit, ExpNZ);
auto ZeroConst = _Constant(0);
Ref SigNZ = _And(OpSize::i64Bit, Gpr, _Constant(0x800f'ffff'ffff'ffffLL));
SigNZ = _Or(OpSize::i64Bit, SigNZ, _Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1);
}
// Comparison and select to push onto stack
SaveNZCV();
_TestNZ(OpSize::i64Bit, Gpr, _Constant(0x7fff'ffff'ffff'ffffUL));
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 3), Size, MEM_OFFSET_SXTX, 1);
}
Ref Sig = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, ExpZV, ExpNZV);
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 6), Size, MEM_OFFSET_SXTX, 1);
}
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
for (int i = 0; i < 7; ++i) {
Ref data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
data = _F80CVTTo(data, 8);
_StoreMem(FPRClass, 16, data, Mem, _Constant((Size * 7) + (i * 10)), 1, MEM_OFFSET_SXTX, 1);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
}
// The final st(7) needs a bit of special handling here
Ref data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
data = _F80CVTTo(data, 8);
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
_StoreMem(FPRClass, 8, data, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
auto topBytes = _VDupElement(16, 2, data, 4);
_StoreMem(FPRClass, 2, topBytes, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
// reset to default
FNINITF64(Op);
_PopStackDestroy();
_PushStack(Exp, Exp, OpSize::i64Bit, true);
_PushStack(Sig, Sig, OpSize::i64Bit, true);
}
// This function converts from F80 on load for compatibility
void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
_StackForceSlow();
const auto Size = GetSrcSize(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = NewFCW;
auto roundShift = _Constant(10);
auto roundMask = _Constant(3);
roundingMode = _Lshr(OpSize::i32Bit, roundingMode, roundShift);
roundingMode = _And(OpSize::i32Bit, roundingMode, roundMask);
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
Ref Top = ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
}
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
auto low = _Constant(~0ULL);
auto high = _Constant(0xFFFF);
Ref Mask = _VCastFromGPR(16, 8, low);
Mask = _VInsGPR(16, 8, 1, Mask, high);
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMem(FPRClass, 16, Mem, _Constant((Size * 7) + (i * 10)), 1, MEM_OFFSET_SXTX, 1);
// Mask off the top bits
Reg = _VAnd(16, 16, Reg, Mask);
// Convert to double precision
Reg = _F80CVT(8, Reg);
_StoreContextIndexed(Reg, Top, 8, MMBaseOffset(), 16, FPRClass);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
}
// The final st(7) needs a bit of special handling here
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
Ref Reg = _LoadMem(FPRClass, 8, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
Ref RegHigh = _LoadMem(FPRClass, 2, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
Reg = _VInsElement(16, 2, 4, 0, Reg, RegHigh);
Reg = _F80CVT(8, Reg); // Convert to double precision
_StoreContextIndexed(Reg, Top, 8, MMBaseOffset(), 16, FPRClass);
}
} // namespace FEXCore::IR
@@ -145,7 +145,7 @@ std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
// These three are all X87 instructions
{0x9B, 1, X86InstInfo{"FWAIT", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9C, 1, X86InstInfo{"PUSHF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF), 0, nullptr}},
{0x9D, 1, X86InstInfo{"POPF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF), 0, nullptr}},
{0x9D, 1, X86InstInfo{"POPF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_BLOCK_END, 0, nullptr}},
{0x9E, 1, X86InstInfo{"SAHF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9F, 1, X86InstInfo{"LAHF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
@@ -21,49 +21,60 @@ constexpr uint16_t PF_3A_66 = 1;
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps = []() consteval {
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct H0F3ATable[] = {
{OPD(0, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(0, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0C), 1, X86InstInfo{"BLENDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0D), 1, X86InstInfo{"BLENDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
auto TableGen = []<uint16_t REX>() consteval {
constexpr U16U8InfoStruct Table[] = {
{OPD(REX, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0C), 1, X86InstInfo{"BLENDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0D), 1, X86InstInfo{"BLENDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x15), 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x15), 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x20), 1, X86InstInfo{"PINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(0, PF_3A_66, 0x21), 1, X86InstInfo{"INSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(0, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x20), 1, X86InstInfo{"PINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x21), 1, X86InstInfo{"INSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x61), 1, X86InstInfo{"PCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x63), 1, X86InstInfo{"PCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x61), 1, X86InstInfo{"PCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x63), 1, X86InstInfo{"PCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_NONE, 0xCC), 1, X86InstInfo{"SHA1RNDS4", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_NONE, 0xCC), 1, X86InstInfo{"SHA1RNDS4", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
};
return std::to_array(Table);
};
constexpr auto H0F3ATable_IgnoresREX0 = TableGen.template operator()<0>();
constexpr auto H0F3ATable_IgnoresREX1 = TableGen.template operator()<1>();
GenerateTable(&Table.at(0), H0F3ATable, std::size(H0F3ATable));
GenerateTable(&Table.at(0), &H0F3ATable_IgnoresREX0.at(0), H0F3ATable_IgnoresREX0.size());
GenerateTable(&Table.at(0), &H0F3ATable_IgnoresREX1.at(0), H0F3ATable_IgnoresREX1.size());
constexpr U16U8InfoStruct TableNeedsREX[] = {
{OPD(0, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
};
GenerateTable(&Table.at(0), TableNeedsREX, std::size(TableNeedsREX));
IR::InstallToTable(Table, IR::OpDispatch_H0F3ATableIgnoreREX);
IR::InstallToTable(Table, IR::OpDispatch_H0F3ATableNeedsREX0);
IR::InstallToTable(Table, IR::OpDispatch_H0F3ATable);
return Table;
}();
void InitializeH0F3ATables(Context::OperatingMode Mode) {
static constexpr U16U8InfoStruct H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRQ", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
};
@@ -67,41 +67,41 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_6, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// GROUP 7
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
// GROUP 8
{OPD(TYPE_GROUP_8, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -23,7 +23,7 @@ auto BaseOpsLambda = []() consteval {
{0x01, 1, X86InstInfo{"", TYPE_GROUP_7, FLAGS_NO_OVERLAY, 0, nullptr}},
// These two load segment register data
{0x02, 1, X86InstInfo{"LAR", TYPE_UNDEC, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x03, 1, X86InstInfo{"LSL", TYPE_UNDEC, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x03, 1, X86InstInfo{"LSL", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0, nullptr}},
{0x04, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x05, 1, X86InstInfo{"SYSCALL", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 0, nullptr}},
{0x06, 1, X86InstInfo{"CLTS", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -225,7 +225,6 @@ enum InstType {
TYPE_SECONDARY_TABLE_PREFIX,
TYPE_X87_TABLE_PREFIX,
TYPE_VEX_TABLE_PREFIX,
TYPE_XOP_TABLE_PREFIX,
TYPE_INST,
TYPE_X87 = TYPE_INST,
TYPE_INVALID,
@@ -466,14 +465,6 @@ constexpr size_t MAX_VEX_TABLE_SIZE = (1 << 13);
// group select (3 bits for now) | ModRM opcode (3 bits)
constexpr size_t MAX_VEX_GROUP_TABLE_SIZE = (1 << 7);
// XOP
// group (2 bits for now) | vex.pp (2 bits) | opcode (8bit)
constexpr size_t MAX_XOP_TABLE_SIZE = (1 << 13);
// XOP group ops
// group select (2 bits for now) | modrm opcode (3 bits)
constexpr size_t MAX_XOP_GROUP_TABLE_SIZE = (1 << 6);
extern std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps;
extern std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps;
extern std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> RepModOps;
@@ -492,10 +483,6 @@ extern std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps;
extern std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps;
extern std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps;
// XOP
extern std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> XOPTableOps;
extern std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> XOPTableGroupOps;
template <typename OpcodeType>
struct X86TablesInfoStruct {
OpcodeType first;
@@ -1,143 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
#include <iterator>
#include <stdint.h>
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> XOPTableOps = []() consteval {
std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> Table{};
#define OPD(group, pp, opcode) ( (group << 10) | (pp << 8) | (opcode))
constexpr uint16_t XOP_GROUP_8 = 0;
constexpr uint16_t XOP_GROUP_9 = 1;
constexpr uint16_t XOP_GROUP_A = 2;
constexpr U16U8InfoStruct XOPTable[] = {
// Group 8
{OPD(XOP_GROUP_8, 0, 0x85), 1, X86InstInfo{"VPMAXSSWW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x86), 1, X86InstInfo{"VPMACSSWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x87), 1, X86InstInfo{"VPMAXSSDQL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x8E), 1, X86InstInfo{"VPMACSSDD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x8F), 1, X86InstInfo{"VPMACSSDQH", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x95), 1, X86InstInfo{"VPMAXSWW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x96), 1, X86InstInfo{"VPMAXSWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x97), 1, X86InstInfo{"VPMAXSDQL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x9E), 1, X86InstInfo{"VPMACSDD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x9F), 1, X86InstInfo{"VPMACSDQH", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xA2), 1, X86InstInfo{"VPCMOV", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xA3), 1, X86InstInfo{"VPPERM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xA6), 1, X86InstInfo{"VPMADCSSWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xB6), 1, X86InstInfo{"VPMADCSWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xC0), 1, X86InstInfo{"VPROTB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xC1), 1, X86InstInfo{"VPROTW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xC2), 1, X86InstInfo{"VPROTD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xC3), 1, X86InstInfo{"VPROTQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xCC), 1, X86InstInfo{"VPCOMccB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xCD), 1, X86InstInfo{"VPCOMccW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xCE), 1, X86InstInfo{"VPCOMccD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xCF), 1, X86InstInfo{"VPCOMccQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xEC), 1, X86InstInfo{"VPCOMccUB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xED), 1, X86InstInfo{"VPCOMccUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xEE), 1, X86InstInfo{"VPCOMccUD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0xEF), 1, X86InstInfo{"VPCOMccUQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
// Group 9
{OPD(XOP_GROUP_9, 0, 0x01), 1, X86InstInfo{"", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, // Group 1
{OPD(XOP_GROUP_9, 0, 0x02), 1, X86InstInfo{"", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, // Group 2
{OPD(XOP_GROUP_9, 0, 0x12), 1, X86InstInfo{"", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, // Group 3
{OPD(XOP_GROUP_9, 0, 0x80), 1, X86InstInfo{"VFRZPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x81), 1, X86InstInfo{"VFRCZPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x82), 1, X86InstInfo{"VFRCZSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x83), 1, X86InstInfo{"VFRCZSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x90), 1, X86InstInfo{"VPROTB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x91), 1, X86InstInfo{"VPROTW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x92), 1, X86InstInfo{"VPROTD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x93), 1, X86InstInfo{"VRPTOQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x94), 1, X86InstInfo{"VPSHLB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x95), 1, X86InstInfo{"VPSHLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x96), 1, X86InstInfo{"VPSHLD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x97), 1, X86InstInfo{"VPSHLQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x98), 1, X86InstInfo{"VPSHAB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x99), 1, X86InstInfo{"VPSHAW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x9A), 1, X86InstInfo{"VPSHAD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0x9B), 1, X86InstInfo{"VPSHAQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xC1), 1, X86InstInfo{"VPHADDBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xC2), 1, X86InstInfo{"VPHADDBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xC3), 1, X86InstInfo{"VPHADDBQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xC6), 1, X86InstInfo{"VPHADDWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xC7), 1, X86InstInfo{"VPHADDWQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xCB), 1, X86InstInfo{"VPHADDDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xD1), 1, X86InstInfo{"VPHADDUBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xD2), 1, X86InstInfo{"VPHADDUBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xD3), 1, X86InstInfo{"VPHADDUBQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xD6), 1, X86InstInfo{"VPHADDUWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xD7), 1, X86InstInfo{"VPHADDUWQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xDB), 1, X86InstInfo{"VPHADDUDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xE1), 1, X86InstInfo{"VPHSUBBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xE2), 1, X86InstInfo{"VPHSUBBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_9, 0, 0xE3), 1, X86InstInfo{"VPHSUBDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
// Group A
{OPD(XOP_GROUP_A, 0, 0x10), 1, X86InstInfo{"BEXTR", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_A, 0, 0x12), 1, X86InstInfo{"", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, // Group 4
};
#undef OPD
GenerateTable(&Table.at(0), XOPTable, std::size(XOPTable));
return Table;
}();
std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> XOPTableGroupOps = []() consteval {
std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> Table{};
#define OPD(subgroup, opcode) (((subgroup - 1) << 3) | (opcode))
constexpr U8U8InfoStruct XOPGroupTable[] = {
// Group 1
{OPD(1, 1), 1, X86InstInfo{"BLCFILL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 2), 1, X86InstInfo{"BLSFILL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 3), 1, X86InstInfo{"BLCS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 4), 1, X86InstInfo{"TZMSK", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 5), 1, X86InstInfo{"BLCIC", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 6), 1, X86InstInfo{"BLSIC", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 7), 1, X86InstInfo{"T1MSKC", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
// Group 2
{OPD(2, 1), 1, X86InstInfo{"BLCMSK", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 6), 1, X86InstInfo{"BLCI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
// Group 3
{OPD(3, 0), 1, X86InstInfo{"LLWPCB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 1), 1, X86InstInfo{"SLWPCB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
// Group 4
{OPD(4, 0), 1, X86InstInfo{"LWPINS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(4, 1), 1, X86InstInfo{"LWPVAL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
};
#undef OPD
GenerateTable(&Table.at(0), XOPGroupTable, std::size(XOPGroupTable));
return Table;
}();
}
+4 -8
View File
@@ -138,7 +138,7 @@ static bool LoadAOTIRCache(AOTIRCacheEntry* Entry, int streamfd) {
auto Array = (AOTIRInlineIndex*)((char*)FilePtr + IndexOffset);
LOGMAN_THROW_AA_FMT(Entry->Array == nullptr && Entry->FilePtr == nullptr, "Entry must not be initialized here");
LOGMAN_THROW_A_FMT(Entry->Array == nullptr && Entry->FilePtr == nullptr, "Entry must not be initialized here");
Entry->Array = Array;
Entry->FilePtr = FilePtr;
Entry->Size = Size;
@@ -338,7 +338,7 @@ bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thre
auto LocalRIP = GuestRIP - AOTIRCacheEntry.VAFileStart;
auto LocalStartAddr = StartAddr - AOTIRCacheEntry.VAFileStart;
auto FileId = AOTIRCacheEntry.Entry->FileId;
const auto& FileId = AOTIRCacheEntry.Entry->FileId;
// The lambda is converted to std::function. This is tricky to refactor so it doesn't allocate memory through glibc.
// NOTE: unique_ptr must be passed as a raw pointer since std::function requires lambda captures to be copyable
@@ -368,10 +368,6 @@ bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thre
}
// Insert to caches if we generated IR
if (GeneratedIR) {
// If the IR doesn't need to be retained then we can just delete it now
delete DebugData;
}
}
return false;
@@ -392,7 +388,7 @@ AOTIRCacheEntry* AOTIRCaptureCache::LoadAOTIRCacheEntry(const fextl::string& fil
auto Inserted = AOTIRCache.insert({fileid, AOTIRCacheEntry {.FileId = fileid, .Filename = filename}});
auto Entry = &(Inserted.first->second);
LOGMAN_THROW_AA_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
LOGMAN_THROW_A_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
if (CTX->Config.AOTIRLoad && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
@@ -409,7 +405,7 @@ AOTIRCacheEntry* AOTIRCaptureCache::LoadAOTIRCacheEntry(const fextl::string& fil
void AOTIRCaptureCache::UnloadAOTIRCacheEntry(AOTIRCacheEntry* Entry) {
#ifndef _WIN32
LOGMAN_THROW_AA_FMT(Entry != nullptr, "Removing not existing entry");
LOGMAN_THROW_A_FMT(Entry != nullptr, "Removing not existing entry");
if (Entry->Array) {
FEXCore::Allocator::munmap(Entry->FilePtr, Entry->Size);
+60 -2
View File
@@ -548,13 +548,16 @@ protected:
// This must directly match bytes to the named opsize.
// Implicit sized IR operations does math to get between sizes.
enum OpSize : uint8_t {
enum class OpSize : uint8_t {
iUnsized = 0,
i8Bit = 1,
i16Bit = 2,
i32Bit = 4,
i64Bit = 8,
f80Bit = 10,
i128Bit = 16,
i256Bit = 32,
iInvalid = 0xFF,
};
enum class FloatCompareOp : uint8_t {
@@ -578,16 +581,71 @@ enum class ShiftType : uint8_t {
// This is a nop operation and will be eliminated by the compiler.
static inline OpSize SizeToOpSize(uint8_t Size) {
switch (Size) {
case 0: return OpSize::iUnsized;
case 1: return OpSize::i8Bit;
case 2: return OpSize::i16Bit;
case 4: return OpSize::i32Bit;
case 8: return OpSize::i64Bit;
case 10: return OpSize::f80Bit;
case 16: return OpSize::i128Bit;
case 32: return OpSize::i256Bit;
case 0xFF: return OpSize::iInvalid;
default: FEX_UNREACHABLE;
}
}
// This is a nop operation and will be eliminated by the compiler.
static inline uint8_t OpSizeToSize(IR::OpSize Size) {
switch (Size) {
case OpSize::iUnsized: return 0;
case OpSize::i8Bit: return 1;
case OpSize::i16Bit: return 2;
case OpSize::i32Bit: return 4;
case OpSize::i64Bit: return 8;
case OpSize::f80Bit: return 10;
case OpSize::i128Bit: return 16;
case OpSize::i256Bit: return 32;
case OpSize::iInvalid: return 0xFF;
default: FEX_UNREACHABLE;
}
}
static inline uint16_t OpSizeAsBits(IR::OpSize Size) {
LOGMAN_THROW_A_FMT(Size != IR::OpSize::iInvalid, "Invalid Size");
return IR::OpSizeToSize(Size) * 8u;
}
template<typename T>
requires (std::is_integral_v<T>)
static inline OpSize operator<<(IR::OpSize Size, T Shift) {
LOGMAN_THROW_A_FMT(Size != IR::OpSize::iInvalid, "Invalid Size");
return IR::SizeToOpSize(IR::OpSizeToSize(Size) << Shift);
}
template<typename T>
requires (std::is_integral_v<T>)
static inline OpSize operator>>(IR::OpSize Size, T Shift) {
LOGMAN_THROW_A_FMT(Size != IR::OpSize::iInvalid, "Invalid Size");
return IR::SizeToOpSize(IR::OpSizeToSize(Size) >> Shift);
}
static inline OpSize operator/(IR::OpSize Size, IR::OpSize Divisor) {
LOGMAN_THROW_A_FMT(Size != IR::OpSize::iInvalid, "Invalid Size");
return IR::SizeToOpSize(IR::OpSizeToSize(Size) / IR::OpSizeToSize(Divisor));
}
template<typename T>
requires (std::is_integral_v<T>)
static inline OpSize operator/(IR::OpSize Size, T Divisor) {
LOGMAN_THROW_A_FMT(Size != IR::OpSize::iInvalid, "Invalid Size");
return IR::SizeToOpSize(IR::OpSizeToSize(Size) / Divisor);
}
static inline uint8_t NumElements(IR::OpSize RegisterSize, IR::OpSize ElementSize) {
LOGMAN_THROW_A_FMT(RegisterSize != IR::OpSize::iInvalid && ElementSize != IR::OpSize::iInvalid, "Invalid Size");
return IR::OpSizeToSize(RegisterSize) / IR::OpSizeToSize(ElementSize);
}
#define IROP_ENUM
#define IROP_STRUCTS
#define IROP_SIZES
@@ -666,7 +724,7 @@ inline NodeID NodeWrapperBase<Type>::ID() const {
bool IsFragmentExit(FEXCore::IR::IROps Op);
bool IsBlockExit(FEXCore::IR::IROps Op);
void Dump(fextl::stringstream* out, const IRListView* IR, IR::RegisterAllocationData* RAData);
void Dump(fextl::stringstream* out, const IRListView* IR, const IR::RegisterAllocationData* RAData);
} // namespace FEXCore::IR
template<>
File diff suppressed because it is too large. Load diff
+36 -19
View File
@@ -82,7 +82,7 @@ static void PrintArg(fextl::stringstream* out, [[maybe_unused]] const IRListView
}
}
static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNodeWrapper Arg, IR::RegisterAllocationData* RAData) {
static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNodeWrapper Arg, const IR::RegisterAllocationData* RAData) {
auto [CodeNode, IROp] = IR->at(Arg)();
const auto ArgID = Arg.ID();
@@ -112,17 +112,17 @@ static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNode
}
if (GetHasDest(IROp->Op)) {
uint32_t ElementSize = IROp->ElementSize;
uint32_t NumElements = IROp->Size;
if (!IROp->ElementSize) {
auto ElementSize = IROp->ElementSize;
uint32_t NumElements = 0;
if (IROp->ElementSize == OpSize::iUnsized) {
ElementSize = IROp->Size;
}
if (ElementSize) {
NumElements /= ElementSize;
if (ElementSize != OpSize::iUnsized) {
NumElements = IR::NumElements(IROp->Size, ElementSize);
}
*out << " i" << std::dec << (ElementSize * 8);
*out << " i" << std::dec << IR::OpSizeAsBits(ElementSize);
if (NumElements > 1) {
*out << "v" << std::dec << NumElements;
@@ -206,6 +206,22 @@ static void PrintArg(fextl::stringstream* out, [[maybe_unused]] const IRListView
return "x87_log10_2";
case NamedVectorConstant::NAMED_VECTOR_X87_LOG_2:
return "x87_log2";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_F32_I32:
return "cvtmax_f32_i32";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_F32_I32_UPPER:
return "cvtmax_f32_i32_upper";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_F32_I64:
return "cvtmax_f32_i64";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_F64_I32:
return "cvtmax_f64_i32";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_F64_I32_UPPER:
return "cvtmax_f64_i32_upper";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_F64_I64:
return "cvtmax_f64_i64";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_I32:
return "cvtmax_i32";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_I64:
return "cvtmax_i64";
default:
return "<Unknown Named Vector Constant>";
}
@@ -221,6 +237,7 @@ static void PrintArg(fextl::stringstream* out, [[maybe_unused]] const IRListView
case OpSize::i64Bit: *out << "i64"; break;
case OpSize::i128Bit: *out << "i128"; break;
case OpSize::i256Bit: *out << "i256"; break;
case OpSize::f80Bit: *out << "f80"; break;
default: *out << "<Unknown OpSize Type>"; break;
}
}
@@ -254,7 +271,7 @@ static void PrintArg(fextl::stringstream* out, [[maybe_unused]] const IRListView
}
}
void Dump(fextl::stringstream* out, const IRListView* IR, IR::RegisterAllocationData* RAData) {
void Dump(fextl::stringstream* out, const IRListView* IR, const IR::RegisterAllocationData* RAData) {
auto HeaderOp = IR->GetHeader();
int8_t CurrentIndent = 0;
@@ -294,14 +311,14 @@ void Dump(fextl::stringstream* out, const IRListView* IR, IR::RegisterAllocation
AddIndent();
if (GetHasDest(IROp->Op)) {
uint32_t ElementSize = IROp->ElementSize;
uint32_t NumElements = IROp->Size;
if (!IROp->ElementSize) {
auto ElementSize = IROp->ElementSize;
uint8_t NumElements = 0;
if (IROp->ElementSize != OpSize::iUnsized) {
ElementSize = IROp->Size;
}
if (ElementSize) {
NumElements /= ElementSize;
if (ElementSize != OpSize::iUnsized) {
NumElements = IR::NumElements(IROp->Size, ElementSize);
}
*out << "%" << std::dec << ID;
@@ -324,7 +341,7 @@ void Dump(fextl::stringstream* out, const IRListView* IR, IR::RegisterAllocation
}
}
*out << " i" << std::dec << (ElementSize * 8);
*out << " i" << std::dec << IR::OpSizeAsBits(ElementSize);
if (NumElements > 1) {
*out << "v" << std::dec << NumElements;
@@ -333,17 +350,17 @@ void Dump(fextl::stringstream* out, const IRListView* IR, IR::RegisterAllocation
*out << " = ";
} else {
uint32_t ElementSize = IROp->ElementSize;
if (!IROp->ElementSize) {
auto ElementSize = IROp->ElementSize;
if (IROp->ElementSize == OpSize::iUnsized) {
ElementSize = IROp->Size;
}
uint32_t NumElements = 0;
if (ElementSize) {
NumElements = IROp->Size / ElementSize;
if (ElementSize != OpSize::iUnsized) {
NumElements = IR::NumElements(IROp->Size, ElementSize);
}
*out << "(%" << std::dec << ID << ' ';
*out << 'i' << std::dec << (ElementSize * 8);
*out << 'i' << std::dec << IR::OpSizeAsBits(ElementSize);
if (NumElements > 1) {
*out << 'v' << std::dec << NumElements;
}
+1 -1
View File
@@ -160,7 +160,7 @@ IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(Ref insertA
if (insertAfter) {
LinkCodeBlocks(insertAfter, CodeNode);
} else {
LOGMAN_THROW_AA_FMT(CurrentCodeBlock != nullptr, "CurrentCodeBlock must not be null here");
LOGMAN_THROW_A_FMT(CurrentCodeBlock != nullptr, "CurrentCodeBlock must not be null here");
// Find last block
auto LastBlock = CurrentCodeBlock;
+19 -15
View File
@@ -11,7 +11,6 @@
#include <FEXCore/fextl/vector.h>
#include <algorithm>
#include <new>
#include <stdint.h>
#include <string.h>
@@ -59,12 +58,12 @@ public:
#define IROP_ALLOCATE_HELPERS
#define IROP_DISPATCH_HELPERS
#include <FEXCore/IR/IRDefines.inc>
IRPair<IROp_Constant> _Constant(uint8_t Size, uint64_t Constant) {
IRPair<IROp_Constant> _Constant(IR::OpSize Size, uint64_t Constant) {
auto Op = AllocateOp<IROp_Constant, IROps::OP_CONSTANT>();
uint64_t Mask = ~0ULL >> (64 - Size);
uint64_t Mask = ~0ULL >> (64 - IR::OpSizeAsBits(Size));
Op.first->Constant = (Constant & Mask);
Op.first->Header.Size = Size / 8;
Op.first->Header.ElementSize = Size / 8;
Op.first->Header.Size = Size;
Op.first->Header.ElementSize = Size;
return Op;
}
IRPair<IROp_Jump> _Jump() {
@@ -77,24 +76,24 @@ public:
return _CondJump(ssa0, _Constant(0), ssa1, ssa2, cond, GetOpSize(ssa0));
}
// TODO: Work to remove this implicit sized Select implementation.
IRPair<IROp_Select> _Select(uint8_t Cond, Ref ssa0, Ref ssa1, Ref ssa2, Ref ssa3, uint8_t CompareSize = 0) {
if (CompareSize == 0) {
CompareSize = std::max<uint8_t>(4, std::max<uint8_t>(GetOpSize(ssa0), GetOpSize(ssa1)));
IRPair<IROp_Select> _Select(uint8_t Cond, Ref ssa0, Ref ssa1, Ref ssa2, Ref ssa3, IR::OpSize CompareSize = OpSize::iUnsized) {
if (CompareSize == OpSize::iUnsized) {
CompareSize = std::max(OpSize::i32Bit, std::max(GetOpSize(ssa0), GetOpSize(ssa1)));
}
return _Select(IR::SizeToOpSize(std::max<uint8_t>(4, std::max<uint8_t>(GetOpSize(ssa2), GetOpSize(ssa3)))),
IR::SizeToOpSize(CompareSize), CondClassType {Cond}, ssa0, ssa1, ssa2, ssa3);
return _Select(std::max(OpSize::i32Bit, std::max(GetOpSize(ssa2), GetOpSize(ssa3))), CompareSize, CondClassType {Cond}, ssa0, ssa1, ssa2, ssa3);
}
IRPair<IROp_LoadMem> _LoadMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref ssa0, uint8_t Align = 1) {
IRPair<IROp_LoadMem> _LoadMem(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref ssa0, IR::OpSize Align = OpSize::i8Bit) {
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
IRPair<IROp_LoadMemTSO> _LoadMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref ssa0, uint8_t Align = 1) {
IRPair<IROp_LoadMemTSO> _LoadMemTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref ssa0, IR::OpSize Align = OpSize::i8Bit) {
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
IRPair<IROp_StoreMem> _StoreMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref Addr, Ref Value, uint8_t Align = 1) {
IRPair<IROp_StoreMem> _StoreMem(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref Addr, Ref Value, IR::OpSize Align = OpSize::i8Bit) {
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
IRPair<IROp_StoreMemTSO> _StoreMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref Addr, Ref Value, uint8_t Align = 1) {
IRPair<IROp_StoreMemTSO>
_StoreMemTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref Addr, Ref Value, IR::OpSize Align = OpSize::i8Bit) {
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
Ref Invalid() {
@@ -206,7 +205,7 @@ public:
ReplaceAllUsesWithRange(Node, NewNode, Start, AllNodesIterator(DualListData.ListBegin(), DualListData.DataBegin()));
LOGMAN_THROW_AA_FMT(Node->NumUses == 0, "Node still used");
LOGMAN_THROW_A_FMT(Node->NumUses == 0, "Node still used");
auto IROp = Node->Op(DualListData.DataBegin())->CW<FEXCore::IR::IROp_Header>();
// We can not remove the op if there are side-effects
@@ -343,8 +342,13 @@ protected:
return Ptr;
}
// MMX State can be either MMX (for 64bit) or x87 FPU (for 80bit)
enum { MMXState_MMX, MMXState_X87 } MMXState = MMXState_MMX;
// Overriden by dispatcher, stubbed for IR tests
virtual void RecordX87Use() {}
virtual void ChgStateX87_MMX() {}
virtual void ChgStateMMX_X87() {}
virtual void SaveNZCV(IROps Op) {}
Ref CurrentWriteCursor = nullptr;
@@ -147,7 +147,6 @@ private:
class IRListView final {
public:
IRListView() = delete;
IRListView(IRListView&&) = delete;
IRListView(DualIntrusiveAllocator* Data)
: IRListView(reinterpret_cast<void*>(Data->DataBegin()), reinterpret_cast<void*>(Data->ListBegin()), Data->DataSize(), Data->ListSize()) {}
+1 -1
View File
@@ -70,7 +70,7 @@ void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl* ctx) {
FEX_CONFIG_OPT(DisablePasses, O0);
if (!DisablePasses()) {
InsertPass(CreateX87StackOptimizationPass());
InsertPass(CreateX87StackOptimizationPass(ctx->HostFeatures));
InsertPass(CreateConstProp(ctx->HostFeatures.SupportsTSOImm9, &ctx->CPUID));
InsertPass(CreateDeadFlagCalculationEliminination());
}
+3 -2
View File
@@ -5,7 +5,8 @@
namespace FEXCore {
class CPUIDEmu;
}
struct HostFeatures;
} // namespace FEXCore
namespace FEXCore::Utils {
class IntrusivePooledAllocator;
@@ -19,7 +20,7 @@ class RegisterAllocationData;
fextl::unique_ptr<FEXCore::IR::Pass> CreateConstProp(bool SupportsTSOImm9, const FEXCore::CPUIDEmu* CPUID);
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass();
fextl::unique_ptr<FEXCore::IR::Pass> CreateX87StackOptimizationPass();
fextl::unique_ptr<FEXCore::IR::Pass> CreateX87StackOptimizationPass(const FEXCore::HostFeatures&);
namespace Validation {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation();
@@ -18,18 +18,13 @@ $end_info$
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/unordered_map.h>
#include <bit>
#include <cstdint>
#include <memory>
#include <optional>
#include <string.h>
#include <tuple>
#include <utility>
namespace FEXCore::IR {
uint64_t getMask(IROp_Header* Op) {
uint64_t NumBits = Op->Size * 8;
uint64_t NumBits = IR::OpSizeAsBits(Op->Size);
return (~0ULL) >> (64 - NumBits);
}
@@ -91,7 +86,7 @@ private:
// We don't allow 8/16-bit operations to have constants, since no
// constant would be in bounds after the JIT's 24/16 shift.
auto Filter = [&IROp](uint64_t X) {
return ARMEmitter::IsImmAddSub(X) && IROp->Size >= 4;
return ARMEmitter::IsImmAddSub(X) && IROp->Size >= OpSize::i32Bit;
};
return InlineIf(IREmit, CurrentIR, CodeNode, IROp, Index, Filter);
@@ -112,7 +107,7 @@ private:
IsSIMM9 &= (SupportsTSOImm9 || !TSO);
// Extended offsets for regular loadstore only.
bool IsExtended = (Imm & (IROp->Size - 1)) == 0 && Imm / IROp->Size <= 4095;
bool IsExtended = (Imm & (IR::OpSizeToSize(IROp->Size) - 1)) == 0 && Imm / IR::OpSizeToSize(IROp->Size) <= 4095;
IsExtended &= !TSO;
if (IsSIMM9 || IsExtended) {
@@ -188,6 +183,35 @@ void ConstProp::HandleConstantPools(IREmitter* IREmit, const IRListView& Current
}
}
// Helper to replace the destination of an instruction with one of its sources,
// to implement algebraic identities. This is surprisingly tricky due to
// implicit masking in our IR.
//
// FEX's IR uses sized opcodes, matching arm64 semantics. 64-bit opcodes do not
// mask, whereas smaller opcodes mask/zero-extend from 32-bits. Therefore, if
// the instruction is 32-bit, we need to mask the source for a sound
// replacement, in case there was garbage in the upper bits.
//
// However, if that source is in turn written by a 32-bit instruction, it is
// guaranteed to have already been masked, so we know there's no garbage and we
// can avoid the zero-extension. This is the case 99% of the time, but the
// masking here is correctness-bearing nevertheless (and new versions of Denuvo
// break if you get this wrong!)
static inline void ReplaceWithSource(IREmitter* IREmit, const IRListView& CurrentIR, Ref CodeNode, IROp_Header* IROp, unsigned Idx) {
Ref Arg = CurrentIR.GetNode(IROp->Args[Idx]);
if (IROp->Size < OpSize::i64Bit) {
LOGMAN_THROW_A_FMT(IROp->Size == OpSize::i32Bit, "other sizes not here");
auto Header = IREmit->GetOpHeader(IROp->Args[Idx]);
if (Header->Size > OpSize::i32Bit) {
Arg = IREmit->_Bfe(OpSize::i32Bit, 32, 0, Arg);
}
}
IREmit->ReplaceAllUsesWith(CodeNode, Arg);
}
// constprop + some more per instruction logic
void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& CurrentIR, Ref CodeNode, IROp_Header* IROp) {
switch (IROp->Op) {
@@ -204,7 +228,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
/* IsImmAddSub assumes the constants are sign-extended, take care of that
* here so we get the optimization for 32-bit adds too.
*/
if (Op->Header.Size == 4) {
if (Op->Header.Size == OpSize::i32Bit) {
Constant1 = (int64_t)(int32_t)Constant1;
Constant2 = (int64_t)(int32_t)Constant2;
}
@@ -285,17 +309,17 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
Replaced = true;
} else if (IROp->Args[0].ID() == IROp->Args[1].ID() || (Constant2 & getMask(IROp)) == getMask(IROp)) {
// AND with same value results in original value
IREmit->ReplaceAllUsesWith(CodeNode, CurrentIR.GetNode(IROp->Args[0]));
ReplaceWithSource(IREmit, CurrentIR, CodeNode, IROp, 0);
Replaced = true;
}
if (!Replaced) {
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IROp->Size * 8); });
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IR::OpSizeAsBits(IROp->Size)); });
}
break;
}
case OP_OR: {
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IROp->Size * 8); });
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IR::OpSizeAsBits(IROp->Size)); });
break;
}
case OP_XOR: {
@@ -318,14 +342,13 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
}
IREmit->SetWriteCursor(CodeNode);
Ref Arg = CurrentIR.GetNode(IROp->Args[1 - i]);
IREmit->ReplaceAllUsesWith(CodeNode, Arg);
ReplaceWithSource(IREmit, CurrentIR, CodeNode, IROp, 1 - i);
Replaced = true;
break;
}
if (!Replaced) {
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IROp->Size * 8); });
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IR::OpSizeAsBits(IROp->Size)); });
}
}
break;
@@ -333,7 +356,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
case OP_ANDWITHFLAGS:
case OP_ANDN:
case OP_TESTNZ: {
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IROp->Size * 8); });
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, [&IROp](uint64_t X) { return IsImmLogical(X, IR::OpSizeAsBits(IROp->Size)); });
break;
}
case OP_NEG: {
@@ -356,13 +379,12 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
if (IREmit->IsValueConstant(IROp->Args[0], &Constant1) && IREmit->IsValueConstant(IROp->Args[1], &Constant2)) {
// Shifts mask the shift amount by 63 or 31 depending on operating size;
uint64_t ShiftMask = IROp->Size == 8 ? 63 : 31;
uint64_t ShiftMask = IROp->Size == OpSize::i64Bit ? 63 : 31;
uint64_t NewConstant = (Constant1 << (Constant2 & ShiftMask)) & getMask(IROp);
IREmit->ReplaceWithConstant(CodeNode, NewConstant);
} else if (IREmit->IsValueConstant(IROp->Args[1], &Constant2) && Constant2 == 0) {
IREmit->SetWriteCursor(CodeNode);
Ref Arg = CurrentIR.GetNode(IROp->Args[0]);
IREmit->ReplaceAllUsesWith(CodeNode, Arg);
ReplaceWithSource(IREmit, CurrentIR, CodeNode, IROp, 0);
} else {
Inline(IREmit, CurrentIR, CodeNode, IROp, 1);
}
@@ -373,8 +395,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
if (IREmit->IsValueConstant(IROp->Args[1], &Constant2) && Constant2 == 0) {
IREmit->SetWriteCursor(CodeNode);
Ref Arg = CurrentIR.GetNode(IROp->Args[0]);
IREmit->ReplaceAllUsesWith(CodeNode, Arg);
ReplaceWithSource(IREmit, CurrentIR, CodeNode, IROp, 0);
} else {
Inline(IREmit, CurrentIR, CodeNode, IROp, 1);
}
@@ -384,7 +405,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
auto Op = IROp->C<IR::IROp_Bfe>();
uint64_t Constant;
if (IROp->Size <= 8 && IREmit->IsValueConstant(Op->Src, &Constant)) {
if (IROp->Size <= OpSize::i64Bit && IREmit->IsValueConstant(Op->Src, &Constant)) {
uint64_t SourceMask = Op->Width == 64 ? ~0ULL : ((1ULL << Op->Width) - 1);
SourceMask <<= Op->lsb;
@@ -400,7 +421,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
if (IREmit->IsValueConstant(Op->Src, &Constant)) {
// SBFE of a constant can be converted to a constant.
uint64_t SourceMask = Op->Width == 64 ? ~0ULL : ((1ULL << Op->Width) - 1);
uint64_t DestSizeInBits = IROp->Size * 8;
uint64_t DestSizeInBits = IR::OpSizeAsBits(IROp->Size);
uint64_t DestMask = DestSizeInBits == 64 ? ~0ULL : ((1ULL << DestSizeInBits) - 1);
SourceMask <<= Op->lsb;
@@ -424,11 +445,11 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
uint64_t NewConstant = SourceMask << Op->lsb;
if (ConstantSrc & 1) {
auto orr = IREmit->_Or(IR::SizeToOpSize(IROp->Size), CurrentIR.GetNode(IROp->Args[0]), IREmit->_Constant(NewConstant));
auto orr = IREmit->_Or(IROp->Size, CurrentIR.GetNode(IROp->Args[0]), IREmit->_Constant(NewConstant));
IREmit->ReplaceAllUsesWith(CodeNode, orr);
} else {
// We are wanting to clear the bitfield.
auto andn = IREmit->_Andn(IR::SizeToOpSize(IROp->Size), CurrentIR.GetNode(IROp->Args[0]), IREmit->_Constant(NewConstant));
auto andn = IREmit->_Andn(IROp->Size, CurrentIR.GetNode(IROp->Args[0]), IREmit->_Constant(NewConstant));
IREmit->ReplaceAllUsesWith(CodeNode, andn);
}
}
@@ -596,7 +617,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
case OP_SELECT: {
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 1, ARMEmitter::IsImmAddSub);
uint64_t AllOnes = IROp->Size == 8 ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
uint64_t AllOnes = IROp->Size == OpSize::i64Bit ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
uint64_t Constant2 {};
uint64_t Constant3 {};
@@ -614,7 +635,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
// We always allow source 1 to be zero, but source 0 can only be a
// special 1/~0 constant if source 1 is 0.
if (InlineIfZero(IREmit, CurrentIR, CodeNode, IROp, 1)) {
uint64_t AllOnes = IROp->Size == 8 ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
uint64_t AllOnes = IROp->Size == OpSize::i64Bit ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
InlineIf(IREmit, CurrentIR, CodeNode, IROp, 0, [&AllOnes](uint64_t X) { return X == 1 || X == AllOnes; });
}
break;
@@ -632,7 +653,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
auto EO = NewRIP->C<IR::IROp_EntrypointOffset>();
IREmit->SetWriteCursor(CurrentIR.GetNode(Op->NewRIP));
IREmit->ReplaceNodeArgument(CodeNode, 0, IREmit->_InlineEntrypointOffset(IR::SizeToOpSize(EO->Header.Size), EO->Offset));
IREmit->ReplaceNodeArgument(CodeNode, 0, IREmit->_InlineEntrypointOffset(EO->Header.Size, EO->Offset));
}
}
break;
@@ -646,6 +667,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
case OP_STOREMEM: {
auto Op = IROp->CW<IR::IROp_StoreMem>();
InlineMemImmediate(IREmit, CurrentIR, CodeNode, IROp, Op->Offset, Op->OffsetType, Op->Offset_Index, Op->OffsetScale, false);
InlineIfZero(IREmit, CurrentIR, CodeNode, IROp, Op->Value_Index);
break;
}
case OP_PREFETCH: {
@@ -661,6 +683,13 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
case OP_STOREMEMTSO: {
auto Op = IROp->CW<IR::IROp_StoreMemTSO>();
InlineMemImmediate(IREmit, CurrentIR, CodeNode, IROp, Op->Offset, Op->OffsetType, Op->Offset_Index, Op->OffsetScale, true);
InlineIfZero(IREmit, CurrentIR, CodeNode, IROp, Op->Value_Index);
break;
}
case OP_STOREMEMPAIR: {
auto Op = IROp->CW<IR::IROp_StoreMemPair>();
InlineIfZero(IREmit, CurrentIR, CodeNode, IROp, Op->Value1_Index);
InlineIfZero(IREmit, CurrentIR, CodeNode, IROp, Op->Value2_Index);
break;
}
case OP_MEMCPY: {
@@ -671,6 +700,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
case OP_MEMSET: {
auto Op = IROp->CW<IR::IROp_MemSet>();
Inline(IREmit, CurrentIR, CodeNode, IROp, Op->Direction_Index);
InlineIfZero(IREmit, CurrentIR, CodeNode, IROp, Op->Value_Index);
break;
}
@@ -27,7 +27,7 @@ private:
};
IRDumper::IRDumper() {
const auto DumpIRStr = DumpIR();
const auto& DumpIRStr = DumpIR();
if (DumpIRStr == "stderr" || DumpIRStr == "stdout" || DumpIRStr == "no") {
// Intentionally do nothing
} else if (DumpIRStr == "server") {
@@ -53,7 +53,7 @@ void IRDumper::Run(IREmitter* IREmit) {
auto IR = IREmit->ViewIR();
auto HeaderOp = IR.GetHeader();
LOGMAN_THROW_AA_FMT(HeaderOp->Header.Op == OP_IRHEADER, "First op wasn't IRHeader");
LOGMAN_THROW_A_FMT(HeaderOp->Header.Op == OP_IRHEADER, "First op wasn't IRHeader");
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpToFile) {
@@ -65,7 +65,7 @@ void IRValidation::Run(IREmitter* IREmit) {
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
if (!EntryBlock) {
EntryBlock = BlockNode;
@@ -79,12 +79,12 @@ void IRValidation::Run(IREmitter* IREmit) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
const auto ID = CurrentIR.GetID(CodeNode);
const uint8_t OpSize = IROp->Size;
const auto OpSize = IROp->Size;
if (GetHasDest(IROp->Op)) {
HadError |= OpSize == 0;
HadError |= OpSize == IR::OpSize::iInvalid;
// Does the op have a destination of size 0?
if (OpSize == 0) {
if (OpSize == IR::OpSize::iInvalid) {
Errors << "%" << ID << ": Had destination but with no size" << std::endl;
}
@@ -191,7 +191,7 @@ unsigned DeadFlagCalculationEliminination::FlagsForCondClassType(CondClassType C
case COND_FLEU:
case COND_FGT: return FLAG_N | FLAG_Z | FLAG_V;
default: LOGMAN_THROW_AA_FMT(false, "unknown cond class type"); return FLAG_NZCV;
default: LOGMAN_THROW_A_FMT(false, "unknown cond class type"); return FLAG_NZCV;
}
}
@@ -319,6 +319,7 @@ constexpr FlagInfo ClassifyConst(IROps Op) {
case OP_STOREAF: return FlagInfo::Pack({.Write = FLAG_A, .CanEliminate = true});
case OP_NZCVSELECT:
case OP_NZCVSELECTV:
case OP_NZCVSELECTINCREMENT:
case OP_NEG:
case OP_CONDJUMP:
@@ -353,6 +354,11 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_NZCVSELECTV: {
auto Op = IROp->CW<IR::IROp_NZCVSelectV>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_NEG: {
auto Op = IROp->CW<IR::IROp_Neg>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
@@ -429,7 +435,7 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
});
}
default: LOGMAN_THROW_AA_FMT(false, "invalid special op"); FEX_UNREACHABLE;
default: LOGMAN_THROW_A_FMT(false, "invalid special op"); FEX_UNREACHABLE;
}
FEX_UNREACHABLE;
@@ -515,7 +521,7 @@ void DeadFlagCalculationEliminination::FoldBranch(IREmitter* IREmit, IRListView&
// Pattern match a branch fed by a compare. We could also handle bit tests
// here, but tbz/tbnz has a limited offset range which we don't have a way to
// deal with yet. Let's hope that's not a big deal.
if (!(Op->Cond == COND_NEQ || Op->Cond == COND_EQ) || (Prev->Size < 4)) {
if (!(Op->Cond == COND_NEQ || Op->Cond == COND_EQ) || (Prev->Size < OpSize::i32Bit)) {
return;
}
@@ -606,7 +612,7 @@ bool DeadFlagCalculationEliminination::ProcessBlock(IREmitter* IREmit, IRListVie
// this flag is outside of the if, since the TestNZ might result from
// optimizing AndWithFlags, and we need to converge locally in a single
// iteration.
if (IROp->Op == OP_TESTNZ && IROp->Size < 4 && !(FlagsRead & (FLAG_N | FLAG_C))) {
if (IROp->Op == OP_TESTNZ && IROp->Size < OpSize::i32Bit && !(FlagsRead & (FLAG_N | FLAG_C))) {
IROp->Op = OP_TESTZ;
}
@@ -21,7 +21,7 @@ using namespace FEXCore;
namespace FEXCore::IR {
namespace {
constexpr uint32_t INVALID_REG = IR::InvalidReg;
[[maybe_unused]] constexpr uint32_t INVALID_REG = IR::InvalidReg;
constexpr uint32_t INVALID_CLASS = IR::InvalidClass.Val;
struct RegisterClass {
@@ -160,7 +160,7 @@ private:
// Otherwise fill from stack
uint32_t SlotPlusOne = SpillSlots[IR->GetID(Old).Value];
LOGMAN_THROW_AA_FMT(SlotPlusOne >= 1, "Old must have been spilled");
LOGMAN_THROW_A_FMT(SlotPlusOne >= 1, "Old must have been spilled");
RegisterClassType RegClass = GetRegClassFromNode(IR, IROp);
@@ -214,7 +214,7 @@ private:
RegisterClass* Class = GetClass(Reg);
uint32_t RegBits = GetRegBits(Reg);
LOGMAN_THROW_AA_FMT(!(Class->Available & RegBits), "Register double-free");
LOGMAN_THROW_A_FMT(!(Class->Available & RegBits), "Register double-free");
Class->Available |= RegBits;
};
@@ -250,7 +250,7 @@ private:
PhysicalRegister DecodeSRAReg(const IROp_Header* IROp, Ref Node) {
RegisterClassType Class;
uint8_t Reg;
uint8_t Reg {};
uint8_t FlagOffset = Classes[GPRFixedClass.Val].Count - 2;
@@ -260,7 +260,7 @@ private:
Class = Op->Class;
Reg = Op->Reg;
} else if (IROp->Op == OP_STOREREGISTER) {
LOGMAN_THROW_AA_FMT(IROp->Op == OP_STOREREGISTER, "node is SRA");
LOGMAN_THROW_A_FMT(IROp->Op == OP_STOREREGISTER, "node is SRA");
const IROp_StoreRegister* Op = IROp->C<IR::IROp_StoreRegister>();
Class = Op->Class;
@@ -289,13 +289,13 @@ private:
// next-use has the /smallest/ unsigned IP.
Ref Candidate = nullptr;
uint32_t BestDistance = UINT32_MAX;
uint8_t BestReg = ~0;
[[maybe_unused]] uint8_t BestReg = ~0;
uint32_t Allocated = ((1u << Class->Count) - 1) & ~Class->Available;
foreach_bit(i, Allocated) {
Ref Old = Class->RegToSSA[i];
LOGMAN_THROW_AA_FMT(Old != nullptr, "Invariant3");
LOGMAN_THROW_A_FMT(Old != nullptr, "Invariant3");
LOGMAN_THROW_A_FMT(SSAToReg[IR->GetID(Map(Old)).Value].Reg == i, "Invariant4");
// Skip any source used by the current instruction, it is unspillable.
@@ -316,11 +316,11 @@ private:
}
}
LOGMAN_THROW_AA_FMT(Candidate != nullptr, "must've found something..");
LOGMAN_THROW_A_FMT(Candidate != nullptr, "must've found something..");
LOGMAN_THROW_A_FMT(IsOld(Candidate), "Invariant5");
PhysicalRegister Reg = SSAToReg[IR->GetID(Map(Candidate)).Value];
LOGMAN_THROW_AA_FMT(Reg.Reg == BestReg, "Invariant6");
LOGMAN_THROW_A_FMT(Reg.Reg == BestReg, "Invariant6");
IROp_Header* Header = IR->GetOp<IROp_Header>(Candidate);
uint32_t Value = IR->GetID(Candidate).Value;
@@ -357,7 +357,7 @@ private:
RegisterClass* Class = GetClass(Reg);
uint32_t RegBits = GetRegBits(Reg);
LOGMAN_THROW_AA_FMT((Class->Available & RegBits) == RegBits, "Precondition");
LOGMAN_THROW_A_FMT((Class->Available & RegBits) == RegBits, "Precondition");
Class->Available &= ~RegBits;
Class->RegToSSA[Reg.Reg] = Unmap(Node);
@@ -435,7 +435,7 @@ private:
}
// Assign a free register in the appropriate class.
LOGMAN_THROW_AA_FMT(Class->Available != 0, "Post-condition of spilling");
LOGMAN_THROW_A_FMT(Class->Available != 0, "Post-condition of spilling");
unsigned Reg = std::countr_zero(Class->Available);
SetReg(CodeNode, PhysicalRegister(ClassType, Reg));
};
@@ -446,7 +446,7 @@ private:
};
void ConstrainedRAPass::AddRegisters(IR::RegisterClassType Class, uint32_t RegisterCount) {
LOGMAN_THROW_AA_FMT(RegisterCount <= INVALID_REG, "Up to {} regs supported", INVALID_REG);
LOGMAN_THROW_A_FMT(RegisterCount <= INVALID_REG, "Up to {} regs supported", INVALID_REG);
Classes[Class].Count = RegisterCount;
}
@@ -623,7 +623,7 @@ void ConstrainedRAPass::Run(IREmitter* IREmit_) {
}
SourceIndex--;
LOGMAN_THROW_AA_FMT(SourceIndex >= 0, "Consistent source count");
LOGMAN_THROW_A_FMT(SourceIndex >= 0, "Consistent source count");
if (!SourcesNextUses[SourceIndex]) {
Ref Old = IR->GetNode(IROp->Args[s]);
@@ -654,11 +654,11 @@ void ConstrainedRAPass::Run(IREmitter* IREmit_) {
}
}
LOGMAN_THROW_AA_FMT(IP >= 1, "IP relative to end of block, iterating forward");
LOGMAN_THROW_A_FMT(IP >= 1, "IP relative to end of block, iterating forward");
--IP;
}
LOGMAN_THROW_AA_FMT(SourceIndex == 0, "Consistent source count in block");
LOGMAN_THROW_A_FMT(SourceIndex == 0, "Consistent source count in block");
}
/* Now that we're done growing things, we can finalize our results.
Loaded 100 of 1320 files, more files were not shown because too many files have changed in this diff. Show more