Compare commits

..
Author SHA1 Message Date
Ryan Houdek ba1b4744c5 Docs: Update for release FEX-2512 2025-12-05 16:11:40 -08:00
Ryan Houdek bd13c02451 Merge pull request #5092 from Sonicadvance1/12
SyscallsSMCTracking: Workaround assert in ELF mapping
2025-12-05 15:44:09 -08:00
Ryan Houdek 5902b175f9 SyscallsSMCTracking: Workaround assert in ELF mapping
The ELF tracking thing has an expectation that only portions of ELF
files that are described in the program headers will be mapped
executable. This doesn't hold true as programs will remap random
portions of ELF files as executable. In the case that this occurs, don't
assert out and instead print a warning.

This was discovered as Node.js remaps a portion of itself executable
that isn't described as such in the program headers. I also have a local
unittest that exposes the same problem. I had discovered same problem in
some other program with #5038.

Also fixes a bug where sometimes completely anonymously mapped
executable sneak in and cause a crash, which is kind of silly.
2025-12-05 14:50:58 -08:00
Ryan Houdek d0e47f9073 Merge pull request #5096 from neobrain/feature_fexofflinecompiler
CodeCache: Implement offline compiler for cache generation
2025-12-05 14:47:58 -08:00
Ryan Houdek bd7215d36f Merge pull request #5093 from Sonicadvance1/13
SteamRT4: Adds support for building the steam depot
2025-12-05 14:39:30 -08:00
Ryan Houdek f3f134f9de FEXServer: Add support for FEX Logging control
Always enables FEXServer log thread when built for Steam so that clients
can be controlled with `STEAM_FEX_LOG=1`. FEX logs will then always go
to FEXServer and those can get directed to wherever pressure-vessel
chooses.
2025-12-04 14:06:18 -08:00
Ryan Houdek cdbf5d57bd Steam: Adds FEXServerManager
This is fairly simple. Needs to be installed alongside FEXServer, so
that it can start it in a portable config.

- Starts a FEXServer
- Tells pressure-vessel when FEXServer is ready
- Keeps FEXServer alive with the `watch_fd` as long as the process lives
- Listens for pressure-vessel to be shutting down
- Exits once pressure-vessel exits, also letting FEXServer shutdown if
  no FEX instances are alive.
2025-12-04 14:06:17 -08:00
Ryan Houdek 1ae50bd670 FEXServerClient: Split out Connect and Start
Allow an optional watch_fd to be passed to FEXServer.
2025-12-04 14:02:43 -08:00
Ryan Houdek d28c9f9843 FEXServerClient: Disallow abstract named sockets under Steam
We don't want clients connecting to random sockets.
2025-12-04 14:02:43 -08:00
Ryan Houdek fe7d52aa78 github/steamrt4: Add artifacts 2025-12-04 14:02:43 -08:00
Ryan Houdek fc0907f8c1 Steam: Add FEXCompatTool 2025-12-04 14:02:43 -08:00
Ryan Houdek e57678d7a6 Config: Move Steam configs into config system 2025-12-04 14:02:43 -08:00
Ryan Houdek 45e594e806 Utils/StringUtils: Add in-place token replace helper 2025-12-04 14:02:43 -08:00
Ryan Houdek 87e7a0effa CMake: Don't install test thunk if tests aren't enabled 2025-12-04 14:02:43 -08:00
Ryan Houdek 4fd1a35b2a CMake: Disable some installs when building for Steam 2025-12-04 14:02:42 -08:00
Ryan Houdek c460cf0678 Merge pull request #5097 from wcampbell-nv/cmdline-map
Remap /proc/pid/cmdline with PR_SET_MM_MAP
2025-12-04 14:02:19 -08:00
Tony Wasserka 983802da61 CodeCache: Add offline compiler for generating caches 2025-12-04 19:16:51 +01:00
Tony Wasserka 49273e0d59 CodeCache: Add workaround for clang-15's broken std::piecewise_construct 2025-12-04 19:16:51 +01:00
Tony Wasserka 76b8459cdc CodeCache: Zero-initialize FileId in ExecutableFileInfo 2025-12-04 19:16:51 +01:00
Tony Wasserka e269eb6f65 ELFCodeLoader: Add helper interfaces 2025-12-04 19:16:51 +01:00
Tony Wasserka 7b4774f375 ELFCodeLoader: Add option to skip interpreter loading 2025-12-04 19:16:51 +01:00
Tony Wasserka 70e9a25112 Syscalls: Move m(un)map to a dedicated interface 2025-12-04 19:16:51 +01:00
Tony Wasserka 9fb83ea56a Windows/CRT: Implement lseek 2025-12-04 19:16:51 +01:00
LC 8bb3398376 Merge pull request #5101 from Sonicadvance1/16
SVE256: Fixes AVX scalar round with insert
2025-12-04 11:24:49 -05:00
Will Campbell d42fbb3d4d Use LoadFileToBuffer + cleanup 2025-12-04 07:54:45 -08:00
Tony Wasserka 53b2245dc1 Merge pull request #5098 from Sonicadvance1/14
docs: Update ProgrammingConcerns
2025-12-04 09:52:59 +00:00
Ryan Houdek db14975828 InstcountCI: Update 2025-12-04 01:45:32 -08:00
Ryan Houdek a5139d2710 unittests/ASM: Adds test case for AVX scalar round with insert bug 2025-12-04 01:44:12 -08:00
Ryan Houdek 9f584c8014 SVE256: Fixes AVX scalar round with insert
We were using the incorrect source registers on SVE256 implementation of
these instructions.

Fixes #5100
2025-12-04 01:43:07 -08:00
Tony Wasserka 6f1b98fb52 docs: Clean up ProgrammingConcerns 2025-12-04 10:18:26 +01:00
Ryan Houdek c258a90505 docs: Update ProgrammingConcerns
Disallow all APIs that touch `FILE`, they all allocate memory that we
don't control.
2025-12-03 20:32:33 -08:00
Will Campbell eb47ef43a7 Read from a fd rather than a FILE 2025-12-03 19:59:13 -08:00
Will Campbell 1876d6b923 Remap cmdline 2025-12-03 18:15:33 -08:00
Ryan Houdek 90c8fcf393 Merge pull request #5094 from pmatos/fix/issue5084
Set current code block in x87 pass
2025-12-02 14:43:54 -08:00
Ryan Houdek 6af90575e9 Merge pull request #5095 from neobrain/feature_serialize_relocations
JIT: Add support for serializing relocations
2025-12-02 14:43:42 -08:00
Tony Wasserka 994613260c CodeCache: Support reverse application of relocations
This allows code to be serialized consistently across runs.
2025-12-02 22:27:53 +01:00
Tony Wasserka 1430fa8220 CodeCache: Move ApplyCodeRelocations 2025-12-02 18:38:59 +01:00
Tony Wasserka 33e06058c6 JIT: Move ApplyRelocations to CodeCache 2025-12-02 18:38:59 +01:00
Tony Wasserka b032d1e1f7 JIT: Make relocations relative to guest base before serialization
This ensures consistency of generated code caches across multiple runs.
2025-12-02 18:38:59 +01:00
Tony Wasserka d6b43b1fe6 Dispatcher: Add public interface to query ExitFunctionLinkerAddress 2025-12-02 17:56:56 +01:00
Paulo Matos fc771c8683 asm_tests: Set current code block in x87 pass 2025-12-02 15:33:08 +01:00
Paulo Matos f91ac09f87 Set current code block in x87 pass
This resets the constant pool in IREmit used by SelectAddressMode().

Fixes #5084.
2025-12-02 15:33:08 +01:00
Ryan Houdek e4fa399412 Merge pull request #5091 from neobrain/feature_better_relocations
JIT: Prepare FEX relocations for code caching
2025-12-01 13:21:29 -08:00
Tony Wasserka 952e949e10 JIT: Clean up block tail writing code 2025-12-01 20:13:51 +01:00
Tony Wasserka 3d093d66fb JIT: Change relocation offset base to CodeBuffer start 2025-12-01 20:13:51 +01:00
Tony Wasserka 05fe2893c7 JIT: Emit relocation from IROP_THUNK 2025-12-01 20:13:51 +01:00
Tony Wasserka 6fc17294b6 JIT: Add relocation for guest RIP stored in jump thunks 2025-12-01 20:13:51 +01:00
Tony Wasserka 6607921bee JIT: Add relocation for guest RIP stored in block tail 2025-12-01 20:13:51 +01:00
Tony Wasserka 0b0793438f JIT: Add relocation for constants relative to the guest entrypoint 2025-12-01 20:13:51 +01:00
LC 3dd591e760 Merge pull request #5088 from Sonicadvance1/11
Linux: Disable io_uring
2025-12-01 14:05:56 -05:00
LC f290d2f899 Merge pull request #5090 from neobrain/refactor_3waycomp
IR: Replace hand-written operators with three-way comparison
2025-12-01 14:05:04 -05:00
Tony Wasserka a676ad7193 JIT: Add explicit padding to relocation descriptors
This ensures zero-initialization, which is required to make code cache
generation produce consistent results.

Also consolidated header fields.
2025-12-01 19:28:39 +01:00
Tony Wasserka 096c408ef6 IR: Replace hand-written operators with three-way comparison 2025-12-01 17:10:06 +01:00
Ryan Houdek 00b65f76b6 Linux: Disable io_uring
This allows passing around `epoll_event` structs which can't be
rewritten due to queues being managed by userspace.
2025-11-30 15:53:34 -08:00
Ryan Houdek 39fb266282 Merge pull request #5087 from neobrain/refactor_simpler_calls
OpcodeDispatcher: Simplify convoluted logic for computing call offsets
2025-11-28 10:10:22 -08:00
Ryan Houdek 3969d0ac78 Merge pull request #5086 from neobrain/fix_invalid_iterators
LookupCache: Fix use of invalidated iterators
2025-11-28 09:55:20 -08:00
Tony Wasserka 439c6bb3c0 OpcodeDispatcher: Simplify convoluted logic for computing call offsets 2025-11-28 11:32:11 +01:00
Tony Wasserka 5cedbf9d34 LookupCache: Fix use of invalidated iterators 2025-11-28 11:29:46 +01:00
Tony Wasserka 427b235eb5 Merge pull request #5071 from Sonicadvance1/2
Github: Add a steamrt4 builder
2025-11-28 08:59:46 +00:00
LC 92d5ba580f Merge pull request #5085 from Sonicadvance1/10
HostFeatures: Extend LRCPC2 errata to more CPUs
2025-11-28 00:41:50 -05:00
Ryan Houdek bc2f331c8b HostFeatures: Extend LRCPC2 errata to C1 Ultra/Premium 2025-11-27 21:29:17 -08:00
Ryan Houdek ca58aef676 HostFeatures: Extend LRCPC2 errata to V3AE 2025-11-27 21:18:51 -08:00
LC 379dc405f6 Merge pull request #5080 from Sonicadvance1/9
FEXInterpreter: Fixes crash with code maps
2025-11-27 20:42:20 -05:00
LC d74b5c42da Merge pull request #5079 from Sonicadvance1/8
FEXServer: Add support for a `wait_fd`
2025-11-27 20:41:22 -05:00
LC 6004971439 Merge pull request #5077 from Sonicadvance1/6
Scripts/InstallFEX: Fixes two issues
2025-11-27 20:39:07 -05:00
LC a12b8927bc Merge pull request #5076 from Sonicadvance1/5
Utils/WritePriorityMutex: Support being forkable
2025-11-27 20:38:37 -05:00
Ryan Houdek 57e23b289a Merge pull request #5081 from esullivan-nvidia/main
HostFeatures: Disable SupportsTSOImm9 for some CPUs
2025-11-27 16:22:40 -08:00
Tony Wasserka 8214ffccf0 Merge pull request #5083 from bylaws/wheiwofjsd
JIT: Fix indirect delinker branch distance
2025-11-27 15:02:47 +00:00
Billy Laws 547135dc2d JIT: Fix indirect delinker branch distance
This is in insts not bytes.
2025-11-27 14:42:56 +00:00
Ryan Houdek 9c72113161 Merge pull request #5070 from wcampbell-nv/cmdline
Reflect application changes to argv[0] in /proc/self/cmdline
2025-11-26 19:50:22 -08:00
esullivan 8cc967fa22 HostFeatures: Disable SupportsTSOImm9 for some CPUs
This change avoids using the LDAPUR instruction with CPUs that are know to be
impacted by an ARM CPU errata that results in poor performance.
2025-11-26 21:42:29 -06:00
Ryan Houdek 4bd30bb72d HostFeatures: Fixes bug in HostFeatures where simulator doesn't support new things
We now have a machine in CI that requires this.
2025-11-26 15:43:37 -08:00
Ryan Houdek 5eeb4dabbd Github: Add a steamrt4 builder
This ensures we don't break downstream projects.
2025-11-26 14:18:08 -08:00
Tony Wasserka a27c4b3860 Merge pull request #5078 from Sonicadvance1/7
CPUID: Fixes regression from #5033
2025-11-26 15:19:51 +00:00
Ryan Houdek dfee08f74f FEXInterpreter: Fixes crash with code maps
When the realpath of a program path can't be resolved, we weren't setting
the config option. This was cascading to be a crashing in codemaps where
it was unconditionally using the optional value (with assert checks),
and causing things to crash.

Pass in the path that can't be resolved to work around a crash in PV
that can happen.
2025-11-25 16:48:47 -08:00
Will Campbell 0e2629bdd4 Address review feedback 2025-11-25 14:09:29 -08:00
Ryan Houdek 05c8630b07 FEXServer: Add support for a wait_fd
This was a requested feature. To make sure that FEXServer is running and
managed by a parent process, we need to have a way to tell FEXServer to
keep alive without any FEX clients. The best way to do this is to pass
FEXServer a Pipe (like FEX does when a client starts it), but instead of
FEXServer signaling to FEXInterpreter that it's ready. FEXServer listens
to the pipe to see if the management process is still alive.

The expectation here is that the management process passes FEXServer the
read end of a pipe, and when the management software is done (or gets
killed by the kernel!) then the write end of the pipe is closed, and
FEXServer naturally closes (As long as there's no FEX processes
remaining).
2025-11-25 12:52:16 -08:00
Ryan Houdek 1cffe618d2 CPUID: Fixes #5033
This leaf changed to being non-constant on that PR since CPUID function
1h returns APICID now.
2025-11-25 11:25:36 -08:00
Ryan Houdek 98c7bb23b5 FEX/InstallFEX: Fixes issue of installing without software-properties-common
Checks to see if the package is installed first before trying to use it.
Fixes an issue where fresh users don't have this package installed and
the script fails.
2025-11-24 14:41:48 -08:00
Ryan Houdek 922853cee1 Scripts/InstallFEX: Fixes #4972
Makes sure that stderr output doesn't cause weird interactions with
FEXRootFSFetcher.
2025-11-24 14:40:56 -08:00
Tony Wasserka a251e61859 Merge pull request #5075 from Sonicadvance1/4
Config: Document the new `FEX_APP_CACHE_LOCATION` option
2025-11-24 20:24:59 +00:00
Ryan Houdek 9c19799023 Config: Document the new FEX_APP_CACHE_LOCATION option
I forgot to document this in the man page.
2025-11-24 12:07:54 -08:00
Ryan Houdek f423b110a8 Utils/WritePriorityMutex: Support being forkable
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.

This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
2025-11-24 11:50:20 -08:00
Ryan Houdek 6fd471e652 Merge pull request #5073 from discapes/unsquashfs-deco-fix
Support detecting unsquashfs>4.7.0 decompressors
2025-11-24 09:05:08 -08:00
Tony Wasserka 3d69029d33 Merge pull request #5072 from Sonicadvance1/3
Minor fixes
2025-11-24 11:24:07 +00:00
Miika Tuominen 2258f2e424 Support detecting unsquashfs>4.7.0 decompressors 2025-11-22 14:38:59 +02:00
Ryan Houdek cca5a68e20 SHMStats: Add missing header 2025-11-21 18:01:37 -08:00
Ryan Houdek da5c9bff68 Async: Add missing header. 2025-11-21 18:01:33 -08:00
Ryan Houdek 5b87f0699b pidof: Switch to using ranges 2025-11-21 18:01:28 -08:00
Will Campbell a67fe561a1 Reflect application changes to argv[0] in /proc/self/cmdline 2025-11-21 16:00:57 -08:00
Tony Wasserka e2f4065376 Merge pull request #5065 from Sonicadvance1/1
FEXCore/CodeCache: Move spin-loop over to a WFE loop
2025-11-21 14:21:45 +01:00
Ryan Houdek d3bf87f4f4 Merge pull request #4985 from Sonicadvance1/fex-atomic
Support (downstream) kernel-side unaligned atomic handling (The rebase sequel)
2025-11-20 17:09:29 -08:00
Ryan Houdek 6d351ec47f FEXCore/CodeCache: Moves spin-loop in to a WFE loop
Saves power and responds faster. Pass in the atomic to `WaitPred` with
the predicate checking if the buffer has been flushed yet. Same
behaviour as previous code but more efficient on our hardware.
2025-11-20 14:23:01 -08:00
Ryan Houdek 4dc1dd2511 FEXCore/Utils/SpinWaitLock: Adds WaitPred for waiting on a predicate
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.

We will need a predicated version for the next commit.
2025-11-20 14:23:01 -08:00
Ryan Houdek 5205ae40fa Merge pull request #5062 from pmatos/fix/address-size-handle
Implement address size modifier handling in CMPSOp and SCASOp
2025-11-20 12:19:22 -08:00
Ryan Houdek 32f1dcde7e Merge pull request #4906 from neobrain/feature_code_maps
CodeCache: Introduce code maps
2025-11-20 12:11:16 -08:00
Ryan Houdek 6772581c53 Arm64ec: Print a log when kernel unaligned atomics are used
Not having this in FEXInterpreter as it is too spammy in the general
case.
2025-11-20 12:05:55 -08:00
Ryan Houdek 387201815b Config: Add option to enable or disable kernel backpatchin on unaligned atomic
Default to enabled because this is the config we expect by default.
In the future will get some benchmarking in various games like
Assassin's Creed, and Call of Duty.
2025-11-20 12:04:59 -08:00
Billy Laws 24d61e1125 Windows: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Billy Laws 04259f031d FEXLoader: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Tony Wasserka 6403da3715 LinuxSyscalls: Shield code map FD from guest access
This prevents chromium/CEF from closing the FD.
2025-11-20 19:13:18 +01:00
Tony Wasserka 90cb76312c LinuxSyscalls: Implement code map writing for future code caching 2025-11-20 19:13:18 +01:00
Tony Wasserka b6cff01abb FEXServer: Add support for querying code maps
Managing code maps in FEXServer rather than in FEXInterpreter makes it
easier to handle multiple concurrent processes sharing code caches for
the main executable and libraries.
2025-11-20 19:13:18 +01:00
Tony Wasserka b34b711161 CodeCache: Add interfaces to describe and generate code maps
Code maps describe per-binary metadata used to generate caches. Currently,
this includes compiled block offsets and loaded shared libraries.
2025-11-20 19:13:18 +01:00
Ryan Houdek 709d767d61 FEX/Config: Allow override of cache location
This will be used.
2025-11-20 18:56:22 +01:00
Tony Wasserka e075916154 Config: Add interface to query cache directory 2025-11-20 18:56:22 +01:00
Paulo Matos de10154f29 instcountci: Implement address size modifier handling in CMPSOp and SCASOp for 64bits 2025-11-20 13:42:19 +01:00
Paulo Matos ba71e79e54 asm_tests: Implement address size modifier handling in CMPSOp and SCASOp 2025-11-20 13:42:19 +01:00
Paulo Matos 2cc70b8051 Implement address size modifier handling in CMPSOp and SCASOp for 64bits
A few games were generating "Can't handle adddress size".
I implemented 0x67 prefix handling for CMPSOp and SCASOP and improved
the error messages for the remainder. This will implement the address
modifier on 64bit systems, and keep issuing an error on 32bits.
2025-11-20 13:42:19 +01:00
Paulo Matos 9d965f94de asm_tests: Add 32-bit CMPS/SCAS tests without address size override 2025-11-20 13:42:19 +01:00
Ryan Houdek 40c2db4744 Merge pull request #5006 from bylaws/fasterrrrr
Introduce two-pass code invalidation model
2025-11-19 17:41:01 -08:00
Billy Laws 9a7285dca4 Windows: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws 8c00ac78b1 Linux: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws cf4478eeee LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-11-20 00:38:03 +00:00
Billy Laws 85c8e7f1bb fextl: Wrap tsl::robin_set 2025-11-20 00:38:03 +00:00
Billy Laws efd95efb40 FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-11-20 00:38:03 +00:00
Billy Laws 99ad7ea45c LookupCache: Drop unused state frame argument for delinker cbs 2025-11-20 00:38:03 +00:00
Ryan Houdek aba0c57f73 Merge pull request #5067 from pmatos/fix/Nasm3
Fix movzx instruction syntax
2025-11-19 14:01:15 -08:00
Paulo Matos 8c4f6b648e Fix movzx instruction syntax
nasm 2.16 was happy with it but it generates a bunch of errors in nasm3.
The generated binaries remain the same.
2025-11-19 15:04:16 +01:00
Ryan Houdek 3b83bdd88d Merge pull request #5066 from neobrain/fix_async_asserts
Async: Adapt precondition checks when receiving FDs
2025-11-19 01:45:41 -08:00
Tony Wasserka 8d71e08b44 Async: Strengthen precondition check when receiving FDs
The sender might provide all requested message bytes but no FD. The receiver
interface has no simple way of indicating this scenario yet, so just assert
out for now to ensure it never happens in the first place.

If needed, this can be changed to return a new error code to indicate partial
read in the future.
2025-11-19 09:48:23 +01:00
Tony Wasserka c31063a8ef Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-19 09:35:52 +01:00
Tony Wasserka 28d101f1dd Revert "Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark"
This cherry-picked an unfinished patch that wasn't intended for merging.
2025-11-19 09:35:08 +01:00
Ryan Houdek aaef344ae3 Merge pull request #5039 from Sonicadvance1/warkwarkwarkwarkwarkwark
FEX: Moves FEX thunk callback function generation to the frontend
2025-11-18 14:40:53 -08:00
Ryan Houdek 42d0324304 FEX: Moves FEX thunk callback function generation to the frontend
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.

Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.

Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.
2025-11-18 14:15:04 -08:00
Ryan Houdek 2a0019347a Thunks: Adds FEX Thunk callback to VDSO
This isn't necessary a VDSO, but it is /always/ mapped in to every
process. Use it as such.
2025-11-18 14:08:43 -08:00
Ryan Houdek e0305ea1b9 Merge pull request #5056 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEX/VDSO: Fixes symbol lookup
2025-11-18 12:51:13 -08:00
Ryan Houdek 105ff47ae3 FEX/VDSO: Fixes symbol lookup
Fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to .dynsym where gcc sticks them in to .symtab. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

Peeled out of #5039
2025-11-18 12:19:51 -08:00
Ryan Houdek 5ee190a41e Merge pull request #5060 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Config: Expose GetConv members
2025-11-17 10:45:02 -08:00
Ryan Houdek cf37617c25 Merge pull request #5011 from pmatos/feat/opt-memcpyf80
Refactoring of storing code in x87 opt. stack pass
2025-11-17 10:25:39 -08:00
Tony Wasserka 2e9c8f0f51 FEXCore/Config: Expose GetConv members
From working branch commit 8244ca1666796267ce25741cdf1103eef4f7539d
`Make cache generation aware of FEX configuration`
2025-11-17 10:21:51 -08:00
Ryan Houdek 06c2319851 Merge pull request #5061 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Common: Adds the ability to override HostFeatures registers by config.
2025-11-17 10:18:26 -08:00
Ryan Houdek da0668c7cc Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
Async: Move file descriptor checks from read_some() to read()
2025-11-17 10:17:44 -08:00
Ryan Houdek b34df334cb Merge pull request #5053 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection
2025-11-17 10:16:59 -08:00
Tony Wasserka 9e9f2ccae1 Merge pull request #5050 from neobrain/refactor_new_config_getter
Config: Refactor value getter interface
2025-11-17 18:58:50 +01:00
Tony Wasserka 15b8f75730 Config: Drop unneeded namespaces from StringArrayType 2025-11-17 18:46:35 +01:00
Tony Wasserka 9d6b9aa574 Config: Allow reading config values in arbitrary C++ expressions
FEX_CONFIG_OPT can only be used as a standalone statement, which is
inconvenient for config values that are only used once. The new functions
(e.g. Get_DUMPIR()) can be used in conditions or other expressions.
2025-11-17 18:46:02 +01:00
Tony Wasserka 64724886af Config: Replace macro-based config readers with a C++ template 2025-11-17 18:46:02 +01:00
Paulo Matos c088369f4a instcountci: Refactoring of storing code in x87 opt. stack pass 2025-11-17 10:14:29 +01:00
Paulo Matos 39dbf46422 Refactoring of storing code in x87 opt. stack pass
Enables memcpy optimization of 80bit floats on reduced precision.

Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
2025-11-17 10:14:29 +01:00
Paulo Matos 3c1b0bb917 Add instcountci tests for 80bit memcpy for x87 instructions 2025-11-17 10:14:29 +01:00
Ryan Houdek 2227170dbb FEXCore/Common: Adds the ability to override HostFeatures registers by config.
This is going to be necessary for the offline compiler work.
Also allow FEXGetConfig to print the same registers in the correct
format for easy fetching.
2025-11-16 16:43:01 -08:00
Ryan Houdek e08f421e1c Convert Async assert to logman assert 2025-11-16 14:56:30 -08:00
Tony Wasserka 42c58c5420 Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-16 14:54:46 -08:00
LC 4afbdd9afb Merge pull request #5055 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fixes alloca use-after-free in `RecvMMsg`
2025-11-14 20:26:05 -05:00
Ryan Houdek 06b9e13904 LinuxSyscalls: Fixes alloca use-after-free in RecvMMsg
Easy enough fix, thanks to @OFFTKP for pointing this out.
2025-11-14 16:27:27 -08:00
Tony Wasserka 29473b43cb LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection 2025-11-14 15:23:42 -08:00
Ryan Houdek 0427d48b98 Merge pull request #5033 from Sonicadvance1/warkwark
CPUID: Fixes APICID for processor count calculation.
2025-11-14 11:31:26 -08:00
Tony Wasserka b228746f1d FEXInterpreter: Fix incorrect value assignment
Previously this was overriding the cached value in the local Value object.
The global SilentLog variable never got updated, so a stale value would be
used.
2025-11-14 11:32:11 +01:00
Tony Wasserka 0be8485116 Merge pull request #5047 from antonkesy/fix_formatting
Align code with clang-format
2025-11-14 09:25:38 +01:00
Tony Wasserka 5d0279ff08 LibraryForwarding/gen: Tiny cleanup 2025-11-14 09:14:45 +01:00
Ryan Houdek 5d908d902c Merge pull request #5045 from lioncash/long
JIT: Handle long ADR/ADRP
2025-11-13 11:10:10 -08:00
Ryan Houdek 3a014f80f2 Merge pull request #5044 from antonkesy/clean_up_scripts
Scripts: Clean-up
2025-11-13 11:09:56 -08:00
Ryan Houdek fbefd7855c Merge pull request #5048 from neobrain/fix_fexconfig_string_lists
FEXConfig: Fix string list handling
2025-11-13 11:07:50 -08:00
Tony Wasserka 11f9135be6 FEXConfig: Fix string list handling 2025-11-13 17:33:21 +01:00
Lioncache 7bb0ce810e JIT: Expand LongAddressGen() to handle movz+movk sequence
This is only ever used on the path where we'd want to handle something
like this (in EmitEntryPoint()), so we can just extend the long handler
type instead of introducing a new type to handle this.
2025-11-13 11:16:34 -05:00
Lioncache 993b832771 JIT: Handle long ADR/ADRP
Wires up the long address handler into the ADR/ADRP restart
handlers.
2025-11-13 09:40:39 -05:00
Anton Kesy 013ac1e627 Align code with clang-format
Automatically done by running:
`find . \( -path './External' -prune \) -o \
  \( -iname '*.cc' -o -iname '*.cpp' -o -iname '*.hpp' -o \
     -iname '*.h' -o -iname '*.c' \) -print | \
  xargs clang-format --style=file -i`
2025-11-13 13:29:49 +01:00
Anton Kesy d7977a02fa remove semicolon 2025-11-12 21:45:09 +01:00
LC 73a32ff22c Merge pull request #5043 from antonkesy/fix_typos
Docs: fix typo
2025-11-12 15:28:24 -05:00
Anton Kesy ee4ae5390b move function comment inside function 2025-11-12 21:22:26 +01:00
Anton Kesy f71db11035 remove excess whitespaces 2025-11-12 21:22:08 +01:00
Anton Kesy 9497288b97 remove unused imports 2025-11-12 21:21:55 +01:00
Anton Kesy a00260d801 remove unused variable 2025-11-12 21:21:28 +01:00
Anton Kesy cad48e07e4 fix comment indentation 2025-11-12 21:21:05 +01:00
Anton Kesy d91e8a4278 docs: fix typo 2025-11-12 21:07:18 +01:00
LC faf74eee90 Merge pull request #5037 from Sonicadvance1/warkwarkwarkwark
FEXCore: Fixes JITGuardPage calculation in a threaded environment
2025-11-12 15:00:00 -05:00
Ryan Houdek 0b52e1cd14 Merge pull request #5042 from neobrain/fix_base_inference
LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
2025-11-12 11:18:58 -08:00
Tony Wasserka b62890f136 LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
At runtime, glxtest is mapped as follows:
0x000055fd9a030000 0x000055fd9a034000 0x4000  0x0     r--p  glxtest
0x000055fd9a034000 0x000055fd9a038000 0x4000  0x3000  r-xp  glxtest
0x000055fd9a038000 0x000055fd9a039000 0x1000  0x6000  rw-p  glxtest
0x000055fd9a039000 0x000055fd9a03a000 0x1000  0x6000  rw-p  glxtest

The problem here is that the last two sections can't be distinguished solely
by their mmap parameters. This would cause the wrong base address to be
inferred for the last mapping. To fix this, we can be more permissive by
allowing multiple candidates to be returned.

In practice, this only affects non-code sections, so it's not a big issue
either way.
2025-11-12 19:54:45 +01:00
Tony Wasserka de1d37eef8 Merge pull request #5041 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwark
CodeEmitter: Removes a few spurious asserts
2025-11-12 14:34:16 +01:00
Ryan Houdek a57c557485 CodeEmitter: Removes a few spurious asserts
These are handled with restart.
2025-11-11 18:25:21 -08:00
LC 3c9f6c845b Merge pull request #5040 from Sonicadvance1/warkwarkwarkwarkwarkwarkwark
FEXCore: Remove usage of "remote atomic" xor
2025-11-11 21:20:18 -05:00
Ryan Houdek ff25e9a92e FEXCore: Remove usage of "remote atomic" xor
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
2025-11-11 17:16:03 -08:00
Ryan Houdek 6a60f72a9e FEXCore: Fixes JITGuardPage calculation in a threaded environment
While this worked great for the singular unit test. I remembered thatour
pool allocator returns the minimum working size asked for but will
return larger sizes if exact fitment couldn't occur.

Because we are dealing with guard pages, we need to return the full
buffer size to the "client" so they can tell the frontend where the
guard page actually lives. Otherwise the JIT will tell the frontend the
guard page is at the end of the requested size, blow past the limit,
and fault in a completely different location.

With a bit of logging I saw in a multithreaded environment that we were
basically always getting a larger requested buffer while Steam was
starting up.
2025-11-11 12:31:38 -08:00
Ryan Houdek 94b690df43 unittests/FEXLinuxTests: Adds cpu core count test to cpuid
Ensures cpuid core counts are reported correctly.
2025-11-11 11:20:26 -08:00
Ryan Houdek 94edbc3436 CPUID: Stop accidentally exposing the HTT bit
We don't support this.
2025-11-11 11:20:26 -08:00
Ryan Houdek 5eab1e559a CPUID: Fixes APICID for processor count calculation.
Primary fix here is returning the current CPU index in function 01h.
Intel Quartus uses this alongside affinity setting to check if all cores
can be used for its calculation. Since we had hardcoded apicid 0 here,
it assumed to only have one core and never generated worker threads.

Additional fix for apicid size. This is the size of the bitmask required
for apic ids, we weren't calculating this correctly at all. This mask is
a "maximum" number of APICs that the CPU reserves in power of two.
Say the core supports 256 APICs, but the processor only supports 16, or
any other combination.
2025-11-11 11:20:26 -08:00
Ryan Houdek 53db3ad6f2 Merge pull request #5036 from neobrain/fix_thunkgen_glibcxx_debug
CMake: Disable libstdc++'s debug mode when compiling thunkgen
2025-11-11 09:14:53 -08:00
Tony Wasserka 581f3263ed CMake: Disable libstdc++'s debug mode when compiling thunkgen
This allows the rest of the project to use _GLIBCXX_DEBUG.
2025-11-11 17:26:37 +01:00
LC 1e3c642be6 Merge pull request #5035 from pmatos/fix/gradual-mem-growth
Use gradual memory growth
2025-11-11 09:00:27 -05:00
LC 22c3cd553f Merge pull request #5034 from Sonicadvance1/warkwarkwark
FEXCore/Win32: Move WritePriorityMutex away from SRWLock
2025-11-11 08:59:29 -05:00
Paulo Matos e5743f8dae Use gradual memory growth
Use min instead of max, otherwise we are always using `MAX_STATS_SIZE`.
2025-11-11 09:23:40 +01:00
Ryan Houdek bddc2f227d FEXCore/Win32: Move WritePriorityMutex away from SRWLock
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.

Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.

While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.

Dark Souls Remastered before:
```
  $RDLck Time: 4.531100 ms/second (0.04 percent)
  $WRLck Time: 2.122560 ms/second (0.02 percent)
```

after:
```
  $RDLck Time: 1.441620 ms/second (0.01 percent)
  $WRLck Time: 0.963720 ms/second (0.01 percent)
```
2025-11-10 17:43:47 -08:00
Ryan Houdek 686c04ea93 Win32: IMplement Wake/Wait by address 2025-11-10 17:43:40 -08:00
Ryan Houdek b38369199e Merge pull request #4893 from Sonicadvance1/long_long_codebuffer_pages
FEX: Implements support for JIT CodeBuffer guard page restart
2025-11-10 13:48:18 -08:00
Ryan Houdek e862c904a9 FEX: Implements support for JIT CodeBuffer guard page restart
When the JIT CodeBuffer overflows, we will now catch accesses to the
guard page and longjump while restarting the JIT with a larger buffer
request.

Fixes #4877
2025-11-10 11:55:21 -08:00
Ryan Houdek 43d9384b1c FEXCore/JIT: Add a pool allocator that understands a guard page
The size asked for has its final page guarded. It's up to the code
asking for allocations to ensure it never uses the final page if
necessary.
2025-11-10 11:54:25 -08:00
Ryan Houdek cb9af0b86a SignalDelegator: Split out SIGSEGV handler
This needs to run before the TestCodeHarness's frontend handler.
2025-11-10 11:53:40 -08:00
Ryan Houdek b8c17a843c ArchHelpers: Adds helper to get pointers to PC and FPRs 2025-11-07 15:33:47 -08:00
Ryan Houdek 7ad7f181d7 FEXCore/LongJump: Add a way to manually load from a longjump
The frontends will need this when loading a longjump buffer in to a
context.
2025-11-07 15:33:47 -08:00
Ryan Houdek eb0bf55033 FEXCore/JIT: Move the JIT long jump buffer to internalthreadstate
This will be a TLS variable that needs to be read by the frontend.
2025-11-07 15:33:47 -08:00
Ryan Houdek f4e3e4ad30 Merge pull request #5031 from lioncash/catch
Externals: Update catch2 from 3.5.3 to 3.11.0
2025-11-07 10:40:49 -08:00
Lioncache 5ae82410cc Externals: Update catch2 from 3.5.3 to 3.11.0
Updates it to the most recent release.
2025-11-07 08:31:04 -05:00
Ryan Houdek 2febb524e9 Merge pull request #5028 from lioncash/fmtup
Externals: Update fmt to 12.1.0
2025-11-06 10:04:15 -08:00
Lioncache b9e452133c Externals: Update fmt to 12.1.0
Keeps fmt updated to its latest release.
2025-11-06 10:17:22 -05:00
LC 747ea0a1f7 Merge pull request #5027 from Sbte/pr/xxhash
Update xxhash to v0.8.3
2025-11-06 07:38:56 -05:00
Sven Baars f8c52ca34a Update xxhash to v0.8.3 2025-11-06 11:53:55 +01:00
Ryan Houdek 663fd5a98b Docs: Update for release FEX-2511 2025-11-05 14:12:52 -08:00
Ryan Houdek 93e58bc15e Merge pull request #5009 from pmatos/feat/stack-xchange-opt
f80 stack xchg optimization for fast path
2025-11-05 12:54:41 -08:00
Ryan Houdek 3ccdf6508e Merge pull request #5024 from neobrain/feature_jit_encoder_recovery
FEXCore/JIT: Add support for recovering from branch encoding failures
2025-11-05 09:27:27 -08:00
Tony Wasserka fd33cf1ce5 Merge pull request #5020 from Sonicadvance1/fix_allocator_bugs
FEXCore/Allocator: Fixes two bugs
2025-11-05 11:53:41 +01:00
Tony Wasserka 2bb64ad1c6 FEXCore/Allocator: Require caller to move unique_ptr into release workaround
This further isolates the workaround to the implementation by highlighting
at the call-site that ownership is moved away.
2025-11-05 11:29:38 +01:00
Paulo Matos e3de62058b instcountci: f80 stack xchg optimization for fast path 2025-11-05 11:20:02 +01:00
Paulo Matos 6ee9984280 f80 stack xchg optimization for fast path 2025-11-05 11:20:02 +01:00
Tony Wasserka e1df548ae9 FEXCore: Extend documentation on uses for UncheckedLongJump 2025-11-05 10:11:30 +01:00
Tony Wasserka 79a685c15e FEXCore: Rename LongJump to UncheckedLongJump
This better reflects the difference to std::longjmp.
2025-11-05 10:11:30 +01:00
Tony Wasserka a61ab2803c CodeEmitter: Drop noisy (un)likely attributes
General usage of these attributes is discouraged. Since this is not
instruction-level performance critical code, drop them.
2025-11-05 09:47:02 +01:00
Ryan Houdek 209ad27332 Code view 2025-11-05 09:47:02 +01:00
Ryan Houdek df08981475 unittests/ASM: Adds test for too large branch objects 2025-11-05 09:47:02 +01:00
Ryan Houdek 5ce6039a02 FEXCore/JIT: Supports restarting JIT in case of encoding failure
ARM64 branches have fairly small relative distances they can encode.
These can be +-1MB, or even +-32KB. The largest relative branch is
+-128MB, which we already set as an upper limit of our block JIT cache
size.

We have for a long time just compiled these without checking with the
expectation that things just happen to work. We didn't hit the asserts
so it was relatively low priority. Apparently now with Steam and a
MaxInst limit of 5000, we are now hitting an assert where we are
encoding too large of a range.

Implement support for long jumping from anywhere in the JIT for when a
long jump tries to be encoded and fails, allowing us to restart the JIT
at any moment. This is implemented as a long jump when this singular
feature could have gotten away with some sort of invasive check and
early exit path for two reasons. For one, that would be even more
invasive, effectively doing try-catch logic manually. And two, the next
step is supporting JIT buffer overflow for when our block size heuristic
fails.

This next step will mandate longjump on SIGSEGV (with cooperative
interaction with the frontend) from effectively /anywhere/ in the JIT.
One of the design goals of the CodeEmitter is that every code emission
function doesn't do a size remaining check to allow the compiler to do
some very effective optimization of emitting code blocks to memory (and
it works!).

But we lose the ability to sanely size check. When writing the emitter I
knew we were going to need to write this cooperative guard page handler,
and we're finally at a point where it needs to be done. This will be in
the next PR although.
2025-11-05 09:47:02 +01:00
Ryan Houdek 6c3fdf723a FEXCore/JIT: Ignore local encoding limit checks
These are guaranteed not to hit encoding distance limits, so we can
ignore the returns.
2025-11-05 09:47:02 +01:00
Ryan Houdek 1a617c1eb2 FEXCore/Dispatcher: Check encoding errors 2025-11-05 09:47:02 +01:00
Ryan Houdek 3a9b801400 FEXCore/VectorRegType: Trivial header fix 2025-11-05 09:47:02 +01:00
Ryan Houdek 6e663acdad Linux/BPFEmitter: Explicitly ignored encoding bool
We know these won't encode in errors.
2025-11-05 09:47:02 +01:00
Ryan Houdek ae1023bb7a unittests/Emitter: Explicitly ignore encoding bool
We know these won't encode in errors.
2025-11-05 09:47:02 +01:00
Ryan Houdek 9cf25e276d CodeEmitter: Return bool if Label instructions can't be encoded
Programming error if they aren't checked, as they will encode
incorrectly if they are too large for their respective instructions.
2025-11-05 09:47:02 +01:00
Ryan Houdek 8db3670ecc FEXCore: Moves longjump implementation from FEX frontend
This will be getting used by FEXCore in a bit.
2025-11-05 09:47:02 +01:00
Ryan Houdek baee367532 FEXCore/Allocator: Move memory leak to a unified location 2025-11-04 16:06:05 -08:00
Ryan Houdek 8da4e72d87 FEXCore/Allocator: Fixes bug where MAP_FIXED could overallocate
When MAP_FIXED is used, if it was larger than the VMA region it was
trying to fit in to, then it would overallocate, corruption memory
adjacent to the VMA region. This was due to a typo in the LiveRegion
range checking.

Fix the typo, add a unittest that tries to overallocate space. Would
assert out without this bug fix.
2025-11-04 16:06:05 -08:00
Ryan Houdek b4a84a2317 Allocator: Fixes false OOM issue in allocator
In the case that overlapping `MAP_FIXED` mmap functions were used, we
were incorrectly tracking the full mapped regions size as new
allocation. We instead need to track which pages have already been
previously allocated and only track those. Would behave like FEX was
running out of memory, but we were just mapping the same location many
times.

Adds a unittest to track this.
2025-11-04 16:06:04 -08:00
Ryan Houdek 43d6347212 FlexBitSet: Add TestAndSet helper 2025-11-04 16:06:04 -08:00
Ryan Houdek 438501e49c Merge pull request #4998 from Sonicadvance1/i_like_my_writes_quick_and_monitored
LookupCache: Convert mutex to new WritePriorityMutex
2025-11-04 16:02:44 -08:00
Ryan Houdek c034e99aaf LookupCache: Convert mutex to new WritePriorityMutex
Changes the single highly-contended lock in `FindBlock` to be a
read-lock.
2025-11-04 15:48:52 -08:00
Ryan Houdek d2d0ca2de9 FEXCore: Implement a write-priority mutex
Now that our Lookup cache mutex is no longer recursive, we can safely
use a shared_mutex instead. The problem with a c++ std::shared_mutex is
that it doesn't guarantee any form of priority, so tens of thousands of
read-locks per second can cause a writer to never acquire the lock, or
take too much time.

The bad news is that C++ doesn't provide us a primitive with
write-priority, so we need to construct our own that is still compatible
with Linux futex. So this is what we do.

- Windows: Uses an SRWLock instead.
  - Only way for WINE to provide us a futex fallback that priorities
    write-priority without stampeding.
2025-11-04 15:48:52 -08:00
LC cbe2b442b2 Merge pull request #5021 from Sonicadvance1/fix_dir_iter
pidof: Fixes another unexpected throw location
2025-11-04 15:10:07 -05:00
LC c5b1cd6e7d Merge pull request #5023 from Sonicadvance1/wark
Wow64: Disable AVX
2025-11-04 15:09:31 -05:00
Ryan Houdek 8d20d1dae3 Wow64: Disable AVX
It's unsupported.
2025-11-04 11:29:04 -08:00
Billy Laws 1f2d702c4e Merge pull request #5004 from pmatos/simp/removeAsFloat
Remove InterpretAsFloat from x87StackOptimizationPass
2025-11-04 10:09:37 +00:00
Ryan Houdek d2d35d0553 pidof: Fixes another unexpected throw location
Exit early if the directory goes away before the iterator is created.
2025-11-03 12:31:03 -08:00
LC 1e7f54dd7e Merge pull request #5019 from Sonicadvance1/fix_flexbitset
FEXCore/Allocator: Fixes FlexBitSet
2025-11-03 08:42:22 -05:00
Ryan Houdek 8430a2f7e6 FEXCore/Allocator: Fixes FlexBitSet
A couple things here, we were never returning the last searched element,
either the last or first depending on search direction.

Also the backward scan would return incorrect indexes in some cases.
Also scanning beyond its page bounds.

Additionally some minorly incorrect assertions.

Adds a new unit test that ensures that we can allocate in to every
location, and that we get the correct indexes back. Also allocated
within guarded pages to ensure it doesn't read outside the bounds.

Fixes a spurious crash in Ender Magnolia.
2025-11-02 18:37:00 -08:00
LC 326cde78e6 Merge pull request #5016 from Sonicadvance1/vulkan_and_gl_fight_tonight
Thunks: Fixes symbol conflict between GL and Vulkan
2025-11-02 13:10:50 -05:00
LC bbc2b0b42f Merge pull request #5017 from Sonicadvance1/describe_esr
ArchHelpers: Adds ESR name helper
2025-11-01 23:20:16 -04:00
Ryan Houdek 231a2c54aa ArchHelpers: Adds ESR name helper
Just helps when an unhandled ESR occurs, it was always a case of needing
to go in to the ARM ARM to decode it which was a bit of a pain. Add a
textual representation of it.
2025-11-01 18:48:24 -07:00
LC 36ae4cee73 Merge pull request #5015 from Sonicadvance1/assert_fix
LookupCache: Fixes assert
2025-11-01 19:52:33 -04:00
LC 62e5ee2201 Merge pull request #5014 from Sonicadvance1/remove_unused_ptrs
FEXCore/CoreState: Removes some unused pointers
2025-11-01 19:51:46 -04:00
Ryan Houdek b89ebd931e Thunks: Fixes symbol conflict between GL and Vulkan
Fixes crash that occurs in applications that use both GL and Vulkan,
Like UE5 Vulkan native games. Fixes Ender Magnolia.

The issue here is that UE5 loads libGL first, which initializes our
libGL thunks, setting its X11Manager's functions.

It then loads libvulkan, which calls our oninit constructor, which
because of the symbol conflict, calls in to the libGL thunk's host
functions to reinitialize its function pointers, never initializing the
Vulkan X11Manager's functions. It would then crash as soon as an X11
function was used.

Give them unique symbol names so we don't accidentally look up the
incorrect symbol.
2025-11-01 16:48:56 -07:00
Ryan Houdek d1b4ddaf61 InstcountCI: Update 2025-11-01 15:18:44 -07:00
Ryan Houdek 734a0b236b FEXCore/CoreState: Removes some unused pointers
NFC
2025-11-01 15:13:10 -07:00
Ryan Houdek 2229c04d4d LookupCache: Fixes assert
These two asserts could never fail, Add assert to the base allocation
instead.
2025-11-01 15:11:44 -07:00
Ryan Houdek 07cff27fa2 SpinWaitLock: Adds one-shot WFE helper 2025-11-01 13:58:22 -07:00
Ryan Houdek 2cf86998bc SpinWaitLock: Fix missing pragma 2025-11-01 13:58:22 -07:00
Ryan Houdek 6ed15a6fd6 Windows: Implement support for SRWLock shared
Exclusive was already implemented.
2025-11-01 13:58:21 -07:00
Ryan Houdek 7610243b0c Merge pull request #5010 from neobrain/fix_asahi_regression
Switch back to jemalloc to fix regression in muvm-based setups
2025-11-01 13:34:36 -07:00
LC e11349b577 Merge pull request #5013 from Sonicadvance1/drm_v6.17
IoctlEmulation: Update to v6.17
2025-11-01 16:17:20 -04:00
Ryan Houdek 9b425697cb IoctlEmulation: Update to v6.17
Nova isn't handled yet because the API is in flux, but it's in v6.17 so
track it.
2025-11-01 13:04:52 -07:00
Ryan Houdek 42596ff91e External/drm-headers: Update to v6.17 2025-11-01 13:04:05 -07:00
LC b3c2ff47f3 Merge pull request #5012 from Sonicadvance1/qemu_apple
FEX: Update CPUID and detect script for newer qemu
2025-11-01 15:56:53 -04:00
Ryan Houdek 75a0bc79be FEX: Update CPUID and detect script for newer qemu
QEmu 10.2 is going to expose MIDR with Apple's vendor ID with variant 0.
That's the best they can do because they don't can't pin threads to
particular cores. So give a string for it, and detect it in the  fit
script.
2025-11-01 12:46:13 -07:00
Tony Wasserka 5a002ad08d Revert "Merge pull request #4969 from Sonicadvance1/rpmalloc"
This reverts commit e1a45a2720, reversing
changes made to bd7edd8651.

The change rendered pressure-vessel non-functional on muvm-based setups
like Fedora Asahi Remix.
2025-10-30 15:22:17 +01:00
Tony Wasserka fbac6f86d1 Merge pull request #5008 from neobrain/fix_removed_option
FEXConfig: Fix crash caused by no longer recognized option
2025-10-29 21:10:46 +01:00
Tony Wasserka ca18bf2a3d FEXConfig: Fix crash caused by no longer recognized option 2025-10-29 20:45:38 +01:00
Ryan Houdek 199effdff7 Merge pull request #5007 from bylaws/fasterefdfdgf
JIT: Restore behaviour of emitting interrupt checks at every block entry
2025-10-28 17:46:18 -07:00
Billy Laws 8212f4b7fb JIT: Restore behaviour of emitting interrupt checks at every block entry
This is needed to handle suspend in infinite loops that occur as a
result of block-size constraints or indirect jumps. Fixes grow home.
2025-10-29 00:34:56 +00:00
Ryan Houdek e1a45a2720 Merge pull request #4969 from Sonicadvance1/rpmalloc
Switch over to rpmalloc instead of jemalloc.
2025-10-28 17:25:25 -07:00
Ryan Houdek bd7edd8651 Merge pull request #5005 from bylaws/oodsakj
Profiler: Fix missing include
2025-10-28 17:25:02 -07:00
Billy Laws 4e1d10a46f Profiler: Fix missing include 2025-10-28 23:52:57 +00:00
Paulo Matos 70b6bc2bae Remove InterpretAsFloat from x87StackOptimizationPass
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.

If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
2025-10-28 11:26:28 +01:00
LC 46d019fe02 Merge pull request #5001 from Sonicadvance1/non_repeating_strings
FEXCore: Have non-repeat strings operations listen to non-tso config
2025-10-27 22:41:37 -04:00
LC d56f689e15 Merge pull request #5002 from Sonicadvance1/fix_typo_in_the_long_long
LinuxSyscalls/Threads: Fixes typo in long jump handler
2025-10-27 22:40:03 -04:00
Ryan Houdek 90702b4102 LinuxSyscalls/Threads: Fixes typo in long jump handler
PR #4892 already found this, but since that isn't merged, make sure this
typo is fixed at least.
2025-10-27 14:05:11 -07:00
Ryan Houdek 0cf105b64a Merge pull request #4993 from Sonicadvance1/moar_stats
FEXCore: Adds some more per-thread stats.
2025-10-27 13:46:12 -07:00
Ryan Houdek 5c74d9458c FEXCore: Have non-repeat strings operations listen to non-tso config
This was missed before, where the non-repeating strings instructions
were still using TSO even when the memcpy/set config option was
disabled. Make sure it listens to the config option and disable TSO in
those instances.

Noticed this while profiling Dishonored, and WINE's `sse2_memmove`
function was showing up as a high amount of CPU time. This is due to
them using non-repeating string operations on the header and tail of
their memmove to align to 16-byte.

With this fixed, it causes the game to go from ~62FPS to ~67FPS,
becoming bottlenecked by x87 emulation instead of memmove. Doing about
23 million soft-float operations per second, because it needs full
precision to remove some flickering artifacts.
2025-10-27 13:05:54 -07:00
Ryan Houdek 75391bf834 External: Remove jemalloc (jemalloc_glibc still exists) 2025-10-27 12:05:08 -07:00
Ryan Houdek 63304a1d88 Windows: rpmalloc 2025-10-27 12:05:08 -07:00
Ryan Houdek 1ab79bd72e FEXCore: Adds some more per-thread stats.
- Cache miss counts
  - Useful for determining if L2 cache or dynamic cache could help
- Cache read/write lock contention times
  - Useful to see if threads are blocking each other on contention
  - Read lock is the case where a read-lock is beneficial, even if we
    currently use a write lock.
- JIT count
  - Useful to see if any new JIT blocks are generating

On top of #4951 because it fiddles with the cache stuff.
2025-10-27 11:25:59 -07:00
Ryan Houdek 985bdf2b6c Switch over to rpmalloc instead of jemalloc.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.

In Bayonetta's title screen it went from 963MB down to 834MB resident.
2025-10-27 11:23:13 -07:00
Ryan Houdek b57ea83aea External: Add rpmalloc 2025-10-27 11:23:13 -07:00
Tony Wasserka d716e22476 Merge pull request #4997 from Sonicadvance1/delete_bad_flags
FEXCore: Remove ABILocalFlags hack
2025-10-27 09:33:35 +01:00
Ryan Houdek bbb8e1ccab FEXCore: Remove ABILocalFlags hack
With our flags being optimized, this does even less than when it was
introduced. It's a hack, people are tinkering with it thinking it'll do
something. Get rid of it.
2025-10-24 17:34:57 -07:00
Ryan Houdek 96f20779c6 Merge pull request #4951 from Sonicadvance1/dynamically_delicious
LookupCache: Adds an option to dynamically scale L1 cache
2025-10-24 14:01:18 -07:00
Ryan Houdek 908313e378 Merge pull request #4996 from cjacek/stlxr-xzr
Arm64: Fix XZR register handling in ARM64EC unaligned STLXR emulation
2025-10-24 13:57:22 -07:00
Jacek Caban 12e5c60633 Arm64: Fix XZR register handling in ARM64EC unaligned STLXR emulation 2025-10-24 22:38:42 +02:00
Ryan Houdek 11946ffc4c InstcountCI: Update 2025-10-24 11:11:54 -07:00
Ryan Houdek f44cd9c545 LookupCache: Adds an option to dynamically scale L1 cache
L1 cache residency can get quite large. Solution, start out small and
scale quickly on L1 cache misses but L2/L3 cache hits.

Some stats on L1 cache residency change:
- Teardown: 40MB -> 16MB (40%)
- Ender Lilies: 79MB -> 32MB (40.5%)
- Death Stranding: 186MB -> 93MB (50%)
- Steam: 75MB -> 7MB (9.3%)

The cost of this option is effectively free in our JIT. It changes a
single LDR to be a single LDP, which on Cortex CPUs cost the same. We do
this by moving the L1 pointer mask in to the CPUState object, making it
dynamic so it lives next to the L1 pointer. We then use that directly
rather than having the hardcoded value.

The lookup cache does a little bit of additional tracking and heuristics
to determine when the current L1 cache should increase or decrease in
size. From 128KB to 16MB per thread, allocating the full VA range as
previously.

Once the heuristic determines that L1 should be increased, it simply
changes the max and the L1 pointer size to compensate, the kernel will
fault in whichever pages are necessary.

Decreasing the size is a little bit more complex, as we want to madvise
the resulting L1 range to ensure we don't have that memory as resident
anymore. Same heuristic but going in the opposite direction otherwise.

Tends to be the case that L1 cache increases a bit on loading screens
then backs down once in-game.

These heuristic values are exposed for increasing and decreasing because
while I think I've picked reasonable values, we will likely need some
more fine tuning over time. Kind of expert user toggles at that point.

Based on #4940 as a base which needs to be merged first.

Full tracked stats from steam as an example of where we are:
```
Total (1000 millisecond sample period):
       JIT Time: 0.486630 ms/second (0.00 percent)
    Signal Time: 0.065880 ms/second (0.00 percent)
     SIGBUS Cnt: 38 (38.160780 per second)
        SMC Cnt: 0
  Softfloat Cnt: 0
FEX JIT Load: 0.004585 (cycles: 552510)
Total FEX Anon memory resident: 368 mB
    JIT resident:             95 mB
    OpDispatcher resident:    38 mB
    Frontend resident:        8 mB
    CPUBackend resident:      624 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 7 mB
    ThreadStates resident:    460 kB
    Unaccounted resident:     217 mB
```
2025-10-24 11:11:54 -07:00
Tony Wasserka 9c5ccb13de Merge pull request #4994 from lioncash/cast
SpinWaitLock, etc: Make use of std::atomic_ref over reinterpret_cast
2025-10-23 17:43:36 +02:00
Lioncache e9bcfd4784 Thread: Make use of std::atomic_ref over cast 2025-10-22 11:07:02 -04:00
Lioncache fd1e8d4566 Arm64: Make use of std::atomic_ref over cast 2025-10-22 11:01:50 -04:00
Lioncache 68dcce0739 SpinWaitLock: Make use of std::atomic_ref over cast
Has a more well-defined way of applying atomic operations to values.
2025-10-22 10:32:49 -04:00
LC 3c554cd787 Merge pull request #4992 from Sonicadvance1/Remove_the_paranoia
FEXCore: Remove Paranoid TSO mode.
2025-10-22 10:05:48 -04:00
LC 9e5f2269d9 Merge pull request #4990 from Sonicadvance1/gettls
wow64/arm64ec: Call GetTLS less frequently
2025-10-21 23:53:36 -04:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
LC eda8ca5449 Merge pull request #4984 from Sonicadvance1/shm_guaranteed_or_your_money_back
SHMStats: Add a 16-byte alignment guarantee
2025-10-21 13:46:53 -04:00
Tony Wasserka 35dd8972e2 Merge pull request #4991 from Sonicadvance1/4096
Removes some hardcoded 4096 constants
2025-10-21 18:28:06 +02:00
Ryan Houdek 643dd56b74 FEX: Removes sone hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:07 -07:00
Ryan Houdek b748eab4ed FEXCore: Removes some hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:06 -07:00
Ryan Houdek f4eaab6977 Merge pull request #4971 from Sonicadvance1/fix_fixington 2025-10-21 06:16:43 -07:00
Ryan Houdek 900c114831 Revert "TestHarnessRunner: Avoid frontend SMC handling"
This reverts commit 2556acb82d.
2025-10-20 16:46:38 -07:00
Ryan Houdek 7937b7e52d unittests: Fixes mixture of code and data in the same page
Test behaviour themselves not changed at all, just data moved or
aligned.

For tests that aren't explicitly testing out SMC behaviour, we were
accidentally relying on some aggressive SMC tracking by mixing data and
code in the same page. To fix this just align the test's data to the
next page boundary which means FEX's SMC tracking won't get triggered
since it is no longer living in the same page.

This has been a thorn for a while, so just get rid of it. We obviously
still have ASM tests that still exist that /do/ rely on SMC, and those
are still expected to work.
2025-10-20 16:46:36 -07:00
Ryan Houdek b40e707771 wow64/arm64ec: Call GetTLS less frequently
If called back-to-back, the compiler can't optimize the object creation
resulting in multiple indirections. Save the creation and pass it
around, allowing the compiler to merge loadstores, and remove redundant
loads.

NFC
2025-10-20 15:45:31 -07:00
Ryan Houdek edde5c8516 Merge pull request #4986 from Sonicadvance1/disable_trace_profiler_default
FEX: Disable trace profiler by default
2025-10-20 12:29:39 -07:00
Ryan Houdek 197facb845 Merge pull request #4980 from neobrain/fix_infer_mapping_base
LinuxSyscalls: Fix base address inference for ELF binaries
2025-10-20 10:26:20 -07:00
Ryan Houdek ca697d0d5d FEX: Disable trace profiler by default
Use a config option to turn it on.
2025-10-20 10:25:17 -07:00
Tony Wasserka 32ddf790d5 LinuxSyscalls: Skip ELF checks for non-header mappings
The full consistency check is only ran in assertion builds now.
2025-10-20 18:00:25 +02:00
Tony Wasserka 2cfba8c6d0 LinuxSyscalls: Fix LookupExecutableFileSection implementation
The file offset of a file mapping doesn't necessarily match its address
offset in virtual memory from the base file mapping. Indeed, most libraries
violate this assumption.

Now that the MappedResource::FirstVMA reliably identifies the base memory
mapping for a given library (even when that library is mapped multiple times),
this can easily be fixed.
2025-10-20 18:00:25 +02:00
Tony Wasserka cf6c0765fa LinuxSyscalls: Create separated MappedResources for re-mappings of the same ELF file
PE/ELF binaries are sometimes mapped multiple times in the same process.
If this happens, there is no longer a unique base virtual address per file.
This breaks assumptions required for code caching: Any time a file mapping
is created, FEX must be able to unambiguously determine the base virtual
address of the mapped library.

This becomes possible by creating a separate MappedResource each time an
ELF header is re-mapped.
2025-10-20 18:00:25 +02:00
Tony Wasserka ad93f27271 Common: Add helper function to find the base virtual address corresponding to an mmap() call 2025-10-20 18:00:25 +02:00
Tony Wasserka 92428e5cbc LinuxSyscalls: Use a multimap to store MappedResources
This allows for creating separate MappedResources when an ELF file is mapped
multiple times at different base addresses.
2025-10-20 18:00:25 +02:00
Tony Wasserka 02c00a87dd LinuxSyscalls: Track MappedResources only for regular files that are executable 2025-10-20 18:00:25 +02:00
Tony Wasserka 6c77bc12c1 LinuxSyscalls: Don't skip MappedResource creation on path query failure
MarkGuestExecutableRange requires a MappedResource to be available even when
FEX doesn't do anything else with it.
2025-10-20 18:00:25 +02:00
LC fbb428c249 Merge pull request #4988 from Sonicadvance1/fix_4982
OpcodeDispatcher: Fix #4982
2025-10-17 23:01:54 -04:00
Ryan Houdek 854a741ea4 OpcodeDispatcher: Fix #4982
Forgot to move the OpcodeDispatcher
2025-10-17 14:01:47 -07:00
Ryan Houdek ed1952a79a Merge pull request #4983 from Sonicadvance1/detect_partial_decode
Frontend: Detect partial decoded instructions
2025-10-17 13:19:27 -07:00
Ryan Houdek 39e8f5122f Merge pull request #4982 from Sonicadvance1/fix_fex_conflict
OpcodeDispatcher: Move FEX reserved instruction
2025-10-17 11:23:09 -07:00
Ryan Houdek 6ee77eb1a1 SHMStats: Adds ThreadStats size to header
Reused the padding area so the header format doesn't change. If it is
non-zero then it should be used by the tool.
2025-10-16 14:07:29 -07:00
Ryan Houdek 46fc45b952 SHMStats: Add a 16-byte alignment guarantee
We want to take advantage of 16-byte single-copy atomicity. Which I am
relying on, but didn't codify it the first time.

Additionally add comments to explain that new members should be added to
the end to allow tools time to gain support gradually. This will allow
me to add new members without fully breaking mangohud, they'll just not
display the new information until support is added.

We're not guaranteeing backwards compatibility, just an attempt not to
constantly churn the format unless necessary. This way if we do break
compatibility, the tool will have an upper bound on supported versions
before needing to rewrite code.
2025-10-16 13:58:31 -07:00
Ryan Houdek 79d90f3c7f unittests: Adds a test to check for partial decode
Known failure so added to the known failure list.
2025-10-16 13:17:16 -07:00
Ryan Houdek eb41cb2261 Frontend: Detect partial decoded instructions
Currently FEX doesn't properly support partial decoded instructions,
which behave slightly differently than full noexec or invalid
instruction decodings. Before this commit we didn't even have a way to
detect the difference.

Primary difference is that the faulting RIP is the beginning of
instruction decode, while the fault address is the first byte that
couldn't be fetched due to memory permissions. This shows up as a
difference between the RIP in mcontext and si_addr in siginfo in the
Linux signal handler.

Right now just change the log so we can determine if we need to support
this edge case.
2025-10-16 13:14:26 -07:00
LC 87e8a9b6aa Merge pull request #4979 from Sonicadvance1/fix_fexpidof_crash
FEXpidof: Fixes potential crash
2025-10-16 14:45:29 -04:00
Ryan Houdek 69f5aa8e35 Merge pull request #4981 from pmatos/revert/nan-work
Revert quiet/signaling nan detection
2025-10-16 11:23:17 -07:00
Ryan Houdek d654f55c3c OpcodeDispatcher: Move FEX reserved instruction
This now conflicts with an SMX instruction, so move it over to another
bytecode that is unlikely to be used.
2025-10-16 11:20:58 -07:00
Paulo Matos eb1689e79f instcountci: Revert Fix quiet and signalling nan propagation 2025-10-16 09:11:01 +02:00
Paulo Matos cf6472df90 Revert "asm_tests: Fix quiet and signalling nan propagation"
This reverts commit c480b0ba41.
2025-10-16 08:53:58 +02:00
Paulo Matos bc6295a78d Revert "Fix quiet and signalling nan propagation"
This reverts commit e7a47a647c.
2025-10-16 08:53:44 +02:00
Ryan Houdek ed5774b88f Merge pull request #4945 from Sonicadvance1/cachy_fexos
FEXCore: Adds option to disable L2 cache lookups
2025-10-15 14:08:14 -07:00
Ryan Houdek 40a29ca9a7 FEXCore: Adds option to disable L2 cache lookups
This saves a whole bunch of memory. Cutting `Just Cause 2`'s title
screen from 1132MB anonymous FEX memory down to 438MB. 629MB in L2
alone.

L2 is primarily a means to reduce overhead in map queries, so it's all
about performance. But because it consumes a lot of people it's kind of
hard.

One idea is that the L2 lookups can be moved to shared data structures,
since we already pull the shared lock when doing an L2 lookup this is
already halfway there.

Side note, we're using unique locks even with read-only code paths
which we can't use the shared lock because this terrible recursive
mutex!

Instead of outright changing L2 behaviour and potentially wrecking
havoc, add a config option for now so testing can happen over time.

before:
```
Total FEX Anon memory resident: 1132 mB
    JIT resident:             60 mB
    OpDispatcher resident:    97 mB
    Frontend resident:        37 mB
    CPUBackend resident:      500 kB
    Lookup cache resident:    629 mB
    Lookup L1 cache resident: 108 mB
    ThreadStates resident:    436 kB
```

after:
```
Total FEX Anon memory resident: 438 mB
    JIT resident:             62 mB
    OpDispatcher resident:    56 mB
    Frontend resident:        22 mB
    CPUBackend resident:      496 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 109 mB
    ThreadStates resident:    436 kB
```
2025-10-15 13:22:19 -07:00
Ryan Houdek a1e1838b11 FEXpidof: Fixes potential crash
Trivial fix, if a symlink gets deleted between checking if it is a
symlink versus getting a path of it then it resulted in a crash.
Happened periodically for me.
2025-10-15 11:41:44 -07:00
Ryan Houdek 4731dbab2f Merge pull request #4965 from Sonicadvance1/remove_recursive_brain_waves
FEXCore: Remove the last recursive_mutex
2025-10-15 09:20:58 -07:00
Ryan Houdek f2841ccb5e FEXCore: Remove the last recursive_mutex
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.

The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.

Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.

I didn't do that exercise since that can be followed up in a subsequent
PR.
2025-10-15 08:42:36 -07:00
Tony Wasserka 81a474fe12 Merge pull request #4978 from dramforever/fix-unittests-llvm-21
unittests/ThunkLibs: Fix build with LLVM 21
2025-10-15 17:36:19 +02:00
dramforever 5af2477d00 unittests/ThunkLibs: Fix build with LLVM 21 2025-10-15 23:04:10 +08:00
Tony Wasserka 4cbba94a29 Merge pull request #4974 from Sonicadvance1/support_deferred_flag
SignalDelegator: Remove bool for deferring signals
2025-10-15 09:15:22 +02:00
Ryan Houdek 2a57428314 Merge pull request #4977 from pmatos/fix/fexserver-socket-uid 2025-10-14 09:07:15 -07:00
Paulo Matos f9ac57205b Fix FEXServer socket resolution
Use getuid() instead of geteuid() when determining FEXServer socket names.
Ensures setuid binaries (like chrome-sandbox from Discord) connect to their parent
user's FEXServer instance instead of trying to spawn a separate server.

Fixes connection errors when running applications that spawn setuid children.
2025-10-14 11:11:24 +02:00
Ryan Houdek 3c6246f99a SignalDelegator: Remove bool for deferring signals
Useful while deferred signals were still a bit more fragile but I
haven't touched this in quite a while no, so just deleted.
2025-10-13 16:47:42 -07:00
Ryan Houdek 072e7bd241 Merge pull request #4973 from pmatos/fix/fcntl64_409
Add F_ADD_SEALS and F_GET_SEALS support to 32-bit fcntl
2025-10-13 11:01:42 -07:00
Ryan Houdek 4080dca816 Merge pull request #4966 from Sonicadvance1/name_remaining
FEX: Name remaining allocations as "Misc"
2025-10-13 10:55:04 -07:00
Ryan Houdek c22fb80129 Merge pull request #4964 from lioncash/vma
SyscallsVMATracking: Make list management internally linked
2025-10-13 10:54:31 -07:00
Paulo Matos 123fcd4bea Add F_ADD_SEALS and F_GET_SEALS support to 32-bit fcntl
Fixes crash with "Unhandled fcntl64: 0x409"
Seen with steam running Bayonetta with Proton Experimental.
2025-10-13 15:27:33 +02:00
Ryan Houdek e0c17672f2 Merge pull request #4970 from bexcran/rdseed-test-fix
unitests/ASM: Fix typo in 09_XX_07.asm
2025-10-11 21:08:39 -07:00
Rebecca Cran 53d484c5f4 unitests/ASM: Fix typo in 09_XX_07.asm
Fix a typo in unittests/ASM/Secondary/09_XX_07.asm which caused it to
jump back to the test_32bit label instead of test_64bit.

On Arm systems with FEAT_RNG support, a significant amount of time may
be required before successive uses of RNDRRS. This is because it returns
a random number with fresh full entropy, and it can take a while to
collect the new entropy. That time may be hundreds or thousands of
instructions, so by jumping back to test_32bit the 64-bit test will
alway fail because an RNDRRS has been executed too recently.

Signed-off-by: Rebecca Cran <rebecca@bsdio.com>
2025-10-11 21:56:41 -06:00
LC a2ea767b77 Merge pull request #4968 from Sonicadvance1/lets_go_questing
InstallFEX:  Add support for 25.10
2025-10-11 14:50:00 -04:00
Ryan Houdek 4bd51a4be0 InstallFEX: Add support for 25.10
The PPA was already updated for this last month.

Fixes #4967
2025-10-11 11:36:32 -07:00
Ryan Houdek 474f2dc267 FEX: Name remaining allocations as "Misc"
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.

With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
        Misc resident:        54 MiB
    JEMalloc resident:        208 MiB
```

So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
2025-10-10 17:34:03 -07:00
Ryan Houdek 033310234e Merge pull request #4963 from lioncash/src
SyscallHandler: Avoid self-referencing global
2025-10-10 12:44:45 -07:00
Lioncache 10835c4b62 SyscallsVMATracking: Make list management internally linked
These don't require being bound to class state directly, and so the list
management can be completely opaque to the outside (also means less
rebuilding if these change)
2025-10-10 15:36:46 -04:00
Lioncache 3e741df77b SyscallHandler: Avoid self-referencing global
We don't need to indirect through the handler global inside of the class
itself for instance member functions.
2025-10-10 15:29:29 -04:00
Ryan Houdek 31d4b4a201 Merge pull request #4960 from Sonicadvance1/enable_avx_32bit
HostFeatures: Enable AVX for 32-bit by default
2025-10-10 10:31:55 -07:00
Ryan Houdek 678985cbf5 Merge pull request #4962 from lioncash/json
JSONPool: Make allocator funcs internally linked
2025-10-09 21:39:17 -07:00
Lioncache e83e9cb6b3 JSONPool: Make allocator funcs internally linked
Makes it a little more explicit that these aren't used outside of this
TU.
2025-10-10 00:11:17 -04:00
Ryan Houdek 48a51679b8 Merge pull request #4961 from lioncash/str
FileManagement: Minor string churn avoiding
2025-10-09 19:29:06 -07:00
Lioncache 987fc925ee FileManagement: Use emplace instead of insert for set
These are cases where the string can be directly constructed internally
instead of being copied/moved.
2025-10-09 22:02:06 -04:00
Lioncache b6cbb23781 FileManagement: Make use of std::string_view in LoadThunkDatabase()
Same behavior, but avoids allocating except for the case where its
necessary.
2025-10-09 22:02:03 -04:00
Ryan Houdek cbb9017c12 Merge pull request #4940 from cjacek/unaligned-ec
ARM64EC: Emulate unaligned atomic access in non-JIT EC code
2025-10-09 15:41:10 -07:00
Ryan Houdek 832d1e8d2c HostFeatures: Enable AVX for 32-bit by default
With the gather overflow fixes in place, I've been having this running
for a while. Now that we just kicked out a release, enable AVX even on
32-bit.

We'll need to eventually create a list of games that explode with AVX
enabled, but that same list would match what happens on real x86 hosts,
so there can be some collaboration there.
2025-10-09 15:18:26 -07:00
Ryan Houdek 34af7f942b unittests/32Bit_ASM: Adds test for VEX.W bug 2025-10-09 15:18:26 -07:00
Ryan Houdek 5f390c16be OpcodeDispatcher: Fixes Scalar FMA size calculation
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.

Correct the implementation to use the correct register size for FMA.
2025-10-09 15:18:26 -07:00
Jacek Caban f36ac0498d Arm64: Emulate LDAXR/STLXR instructions in non-JIT ARM64EC code 2025-10-10 00:14:48 +02:00
Jacek Caban 18bdd9b665 Arm64: Factor out DoCAS 2025-10-10 00:13:57 +02:00
Jacek Caban 5fd3852fd3 ARM64EC: Emulate unaligned atomic access in non-JIT EC code 2025-10-10 00:13:56 +02:00
Ryan Houdek 764432c77c Merge pull request #4959 from lioncash/config
Config: Prevent duplicate lookup in EnvLoader::Load()
2025-10-09 12:09:26 -07:00
Lioncache 55c43229ca Config: Remove unnecessary check in GetConfigDirectory()
This is already checked for in the outer if.
2025-10-09 14:57:00 -04:00
Lioncache dbc4daaaa4 Config: std::move string into EnvMap
Same behavior, but without a reallocation.
2025-10-09 14:57:00 -04:00
Lioncache acdbbec78d Config: Prevent duplicate lookup in EnvLoader::Load()
We can check the iterator itself instead of doing the lookup again.
2025-10-09 14:56:57 -04:00
Ryan Houdek 3247e777c0 Merge pull request #4958 from lioncash/move
Core: Add missing std::move in AddForceTSOInformation()
2025-10-09 09:48:44 -07:00
Lioncache c6e60ff3f5 Core: Add missing std::move in AddForceTSOInformation()
All callsites move the instructions into the function, but we weren't
further passing the rvalue-reference to merge().
2025-10-09 12:36:44 -04:00
Ryan Houdek b7df102919 Merge pull request #4957 from lioncash/bitwise
OpcodeDispatcher: Amend wonky bitwise AND usage in LoadMemPairAutoTSO/_StoreMemPairAutoTSO
2025-10-09 09:03:41 -07:00
Lioncache f5d450b95c OpcodeDispatcher: Amend wonky bitwise AND usage in LoadMemPairAutoTSO/_StoreMemPairAutoTSO 2025-10-09 11:39:17 -04:00
Ryan Houdek 972aaf9f5f Merge pull request #4956 from lioncash/init 2025-10-09 07:49:23 -07:00
Lioncache b98d5f30c0 RegisterAllocationPass: Ensure relevant members are initialized
Ensures that they have deterministic values on construction
2025-10-09 04:17:21 -04:00
Ryan Houdek d2e2e189d8 Merge pull request #4955 from lioncash/bpf
BPFEmitter: Minor cleanup
2025-10-08 22:27:41 -07:00
Lioncache 17d110575c BPFEmitter: Remove unused SeccopEmulator include 2025-10-09 00:57:49 -04:00
Lioncache a967976420 BPFEmitter: Simplify iterator declarations
Makes the jump labels a little quicker to read.
2025-10-09 00:57:49 -04:00
Lioncache 7e1e1ccccc BPFEmitter: Move private API into implementation
Keeps some internals fully private (and also avoids dumping
some defines into other headers).
2025-10-09 00:57:49 -04:00
Ryan Houdek 502452602e Merge pull request #4954 from lioncash/parser
ELFParser: Avoid unnecessary copies
2025-10-08 21:42:20 -07:00
LC f8ff46f3e3 Merge pull request #4952 from Sonicadvance1/naming_block_links
FEXCore/fexl: Support a named monotonic_buffer_resource
2025-10-08 23:04:41 -04:00
LC 8f572deaed Merge pull request #4953 from Sonicadvance1/remove_inlinesyscall
FEXCore: Remove InlineSyscall
2025-10-08 23:02:40 -04:00
Ryan Houdek 601dd1d2b5 InstcountCI: Update 2025-10-08 19:33:58 -07:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek f1f81f9de2 FEXCore: Remove InlineSyscall
Due to IR changes we can no longer do this, its use was fairly limited
anyway.
2025-10-08 18:47:18 -07:00
Lioncache 8513c02ec1 ELFParser: Avoid unnecessary copies
Gets rid of a few 56 byte copies.
2025-10-08 21:20:46 -04:00
Ryan Houdek 320c5f1847 Docs: Update for release FEX-2510 2025-10-08 17:03:06 -07:00
Ryan Houdek 282f091e85 FEXCore/fexl: Support a named monotonic_buffer_resource
Lets us track our memory usage of our block links.
2025-10-08 16:43:28 -07:00
Ryan Houdek 8647033029 FEXCore: Support naming a bunch of VMA regions
Useful for memory usage tracking.
2025-10-08 16:43:28 -07:00
Ryan Houdek 3b88947cbe Merge pull request #4950 from lioncash/syscall
SyscallHandler: Shrink definition struct from 24 to 16 bytes
2025-10-08 14:03:57 -07:00
Ryan Houdek f05be0120e Merge pull request #4949 from lioncash/thread
ThreadManager: Make StatAlloc instance private
2025-10-08 13:44:43 -07:00
Lioncache 10c3b2fb5d SyscallHandler: Shrink definition struct from 24 to 16 bytes
Just a minor space saving, given how many syscall definitions we'll
have.
2025-10-08 16:40:20 -04:00
Lioncache 74ae89865c SHMStats: Add missing virtual destructor
This should be present for universally consistent behavior, regardless
of how the class is used.
2025-10-08 16:17:42 -04:00
Lioncache a2445904e7 ThreadManager: Make StatAlloc instance private
This doesn't need to be public.
2025-10-08 16:13:23 -04:00
Ryan Houdek bd5188e7bc Merge pull request #4948 from lioncash/fmt2
EnumUtils: Remove fmt include
2025-10-07 13:04:22 -07:00
Lioncache d89c119d9d EnumUtils: Remove fmt include
Forgot to remove this in the previous PR.
2025-10-07 12:54:32 -04:00
Ryan Houdek 802f3bed47 Merge pull request #4947 from lioncash/format
EnumUtils: Further simplify enum passthrough formatting
2025-10-07 06:47:56 -07:00
Ryan Houdek 9c606a278d Merge pull request #4946 from lioncash/sys
SysCalls: Remove unimplemented prototypes
2025-10-07 06:46:59 -07:00
Lioncache 7aa5bc0503 EnumUtils: Further simplify enum passthrough formatting
Turns out a simpler way was added to the docs at some point and I never
noticed.

Before:
   text     data      bss      dec      hex  filename
4159895  1471360  4336824  9968079   9819cf  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4157159  1471360  4336824  9965343   980f1f  Bin/FEX
2025-10-07 01:52:05 -04:00
Lioncache d94b9fae94 SourceCodeResolver: Pass string_view by value for GenerateMap()
Generally this should be passed as a value type unless there's a good
reason not to.
2025-10-07 01:27:15 -04:00
Lioncache 05fcaaa758 Syscalls: Remove unimplemented prototypes 2025-10-07 01:27:11 -04:00
Ryan Houdek 9626a64340 Merge pull request #4944 from lioncash/type
Addressing: Shave 8 bytes off AddressMode
2025-10-06 14:01:09 -07:00
Lioncache c1cfd4db83 Addressing: Shave 8 bytes off AddressMode
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.

Before:
   text     data      bss      dec      hex  filename
4160559  1471360  4336824  9968743   981c67  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4159927  1471360  4336824  9968111   9819ef  Bin/FEX
2025-10-06 15:52:15 -04:00
Ryan Houdek bfda91ac16 Merge pull request #4943 from lioncash/table
X86Tables: Remove unused LateInitCopyTable()
2025-10-06 12:17:39 -07:00
Lioncache b7117b86ea X86Tables: Remove unused LateInitCopyTable() 2025-10-06 15:04:36 -04:00
LC a94059ad81 Merge pull request #4942 from Sonicadvance1/fix_map
ThreadManager: Fix mmap usage to not take guest space
2025-10-06 14:59:27 -04:00
Ryan Houdek 2cab87fbd0 ThreadManager: Fix mmap usage to not take guest space
SHM stats and CallRet stack was accidentally using system mmap which
means it would take VA space from the guest. Ensure it uses the
Allocator helpers so that doesn't happen.

Gives 32-bit guests 8-ish megabyte of VA space back.
2025-10-06 11:20:00 -07:00
Ryan Houdek 2017100ff3 Merge pull request #4941 from lioncash/unused
OpcodeDispatcher: Remove unimplemented function prototypes
2025-10-06 10:18:18 -07:00
Lioncache 72cb29d36b OpcodeDispatcher: Remove unimplemented function prototypes
Just cleans out the interface a little.
2025-10-06 13:08:07 -04:00
Ryan Houdek a2b0d594fb Merge pull request #4939 from lioncash/invalid
OpcodeDispatcher: Default alignment parameters for store helpers
2025-10-06 09:11:49 -07:00
Lioncache a16d4ff1f3 OpcodeDispatcher: Deduplicate in LEAOp/SMSWOp
We can shorten a few lines here by just storing the op addr value
to a local variable.
2025-10-06 11:13:00 -04:00
Lioncache 305b1ecf2d OpcodeDispatcher: Default alignment parameters for store helpers
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.

Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
2025-10-06 11:04:12 -04:00
Ryan Houdek 3cc0cae249 Merge pull request #4938 from lioncash/prctl
PrctlUtils: Move to include folder
2025-10-05 23:27:38 -07:00
Lioncache 137aa59254 PrctlUtils: Move to include folder
We can group more prctl value handling in here.
2025-10-06 02:12:29 -04:00
LC 5f8cb0dc54 Merge pull request #4937 from Sonicadvance1/name_vma
Allocator: Name FEX's VMA regions for allocation
2025-10-05 22:40:58 -04:00
Ryan Houdek 45978474f3 Allocator: Name FEX's VMA regions for allocation
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.

Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.

This tracking will be the first step towards seeing where our memory
usage is going.
2025-10-05 19:30:59 -07:00
Ryan Houdek 75206c7a51 Merge pull request #4936 from lioncash/str
Common/StringConv: std::stoull -> std::strtoull for enum handler
2025-10-05 13:14:02 -07:00
Ryan Houdek 06ff0a45a7 Merge pull request #4935 from lioncash/move
IR: Remove Swap1/Swap2 ops
2025-10-05 11:57:35 -07:00
Ryan Houdek c948d532a7 Merge pull request #4934 from lioncash/select
OpcodeDispatcher: Move off implicit _Select
2025-10-05 11:54:59 -07:00
Ryan Houdek 3825483f79 Merge pull request #4933 from lioncash/reg
RedundantFlagCalculationElimination: Eliminate unnecessary vector copy
2025-10-05 11:53:56 -07:00
Lioncache 0b238d942e Common/StringConv: std::stoull -> std::strtoull for enum handler
Missed this when simplifying the conversion facilities, but
we should be using std::strtoull here, as per the programming
concerns.
2025-10-05 14:49:07 -04:00
Lioncache 16b09a9dd3 OpcodeDispatcher: Move off implicit _Select
Resolves a lingering TODO.
2025-10-05 13:57:36 -04:00
Lioncache 4f6800b768 IR: Remove Swap1/Swap2 ops
These are no longer used.
2025-10-05 13:26:51 -04:00
Lioncache c84801adb3 RedundantFlagCalc: Avoid vector copy in OptimizeParity()
Previously this was making a copy of the vector, when we only
need to read from it.
2025-10-05 12:17:45 -04:00
Lioncache 69f990c1f8 RedundantFlagCalc: Organize headers
Also remove incorrect comment that this pass isn't used.
2025-10-05 12:17:45 -04:00
Lioncache 8a83564678 RedundantFlagCalc: Remove unimplemented prototype
Just tidies the interface a little.
2025-10-05 12:17:45 -04:00
Lioncache 43d93f836f RedundantFlagCalc: Mark members as const where applicable
These don't affect member state.
2025-10-05 03:14:21 -04:00
Ryan Houdek 7c2c0f7fe5 Merge pull request #4932 from lioncash/python
json_ir_generator: Minor cleanup
2025-10-04 11:44:08 -07:00
Lioncache 3ff47cf677 json_ir_generator: Add missing error message
ExitError() expects a string argument. Unlikely error case to be hit,
but we should still make the error useful to the reader.
2025-10-04 11:00:05 -04:00
Lioncache 3165dce89e json_ir_generator: Turn IROpNameMap into a set
This is only used to track for duplicates, so we only really need
the key name, since the values are never used.
2025-10-04 11:00:00 -04:00
Lioncache d4fad0a3b0 json_ir_generator: Simplify some looping
We can use enumerate to collapse a few of these
2025-10-04 10:37:34 -04:00
Lioncache 15fcaa794f json_ir_generator: Simplify is_ssa_type
Same behavior, but a little more straightforward
2025-10-04 08:47:08 -04:00
Lioncache 6a1a505336 json_ir_generator: Add type annotations for top level vars
Makes the types a little more explicit and also makes member
suggestions in IDEs (vscode, etc) work a little better.
2025-10-04 08:45:08 -04:00
LC 85937a7bfb Merge pull request #4931 from Sonicadvance1/sse4a
Implement remaining SSE4a instructions
2025-10-04 08:22:58 -04:00
Ryan Houdek 3baa598b9e HostFeatures: Adds flag for SSE4a 2025-10-04 02:51:28 -07:00
Ryan Houdek 6f83b10af6 InstcountCI: Update 2025-10-04 02:46:27 -07:00
Ryan Houdek f52bcb49ab unittests/ASM: Adds new variable SSE4a tests 2025-10-04 02:46:14 -07:00
Ryan Houdek 9cef8ff7ce OpcodeDispatcher: Implement support for SSE4a variable extrq/insertq 2025-10-04 02:44:56 -07:00
Ryan Houdek a499ad6404 IR: Implement two new IR ops
rbit is useful in generic algorithms.
`MaskGenerateFromBitWidth` is only really useful for SSE4a, but
implemented in the OpcodeDispatcher is rough, so add an operation.
2025-10-04 02:43:39 -07:00
Ryan Houdek 7c8767ab32 OpcodeDispatcher: Minor improvement to inserting constant to vector 2025-10-04 02:43:21 -07:00
Ryan Houdek 38a8559943 Merge pull request #4930 from lioncash/forward
IR: Move RegisterAllocationPass forward decl to JITClass
2025-10-03 21:22:17 -07:00
Lioncache 48c6acbc87 IR: Move RegisterAllocationPass forward decl to JITClass
This isn't actually used anywhere in the IR header, so we can
move it to where it's actually used.

Now the IR interface header doesn't have anything related to the
independent passes in it.
2025-10-04 00:08:58 -04:00
Ryan Houdek b5d93bc4a0 Merge pull request #4928 from lioncash/passthru
EnumUtils: Add define for default passthrough formatting
2025-10-03 19:45:41 -07:00
LC a9fffe771b Merge pull request #4929 from Sonicadvance1/sse4a_imm
OpcodeDispatcher: Implement support for imm variants of SSE4a extrq/insertq
2025-10-03 17:40:19 -04:00
Ryan Houdek 69f3ab14b9 InstcountCI: Update 2025-10-03 13:51:51 -07:00
Ryan Houdek a7b9a34ec9 unittests: Adds imm extrq/insertq tests 2025-10-03 13:34:32 -07:00
Ryan Houdek f45be1f59e OpcodeDispatcher: Implement support for imm extrq/insertq
Fairly straightforward to implement, but not exciting in their
performance.
2025-10-03 13:05:43 -07:00
Ryan Houdek 012c2ba851 IR: Fixes vector 64-bit binops
We were only supporting operating size of 256-bit and 128-bit. These
also support 64-bit which wasn't wired up.

SSE4a will want to use 64-bit.
2025-10-03 13:04:09 -07:00
Lioncache a0c2ce0068 EnumUtils: Add define for default passthrough formatting
Handles a normal case where printing an enum type as an integral value
is still desirable.

Mainly just a way to reduce boilerplate.
2025-10-03 14:21:19 -04:00
Ryan Houdek 86898035da Merge pull request #4926 from lioncash/enum
IR: Convert RegisterClassType to an enum class
2025-10-03 10:35:23 -07:00
Ryan Houdek 05eeffe0f7 Merge pull request #4927 from bylaws/ffds
Windows: Avoid redeclaring strtoll as unimplemented
2025-10-03 10:32:39 -07:00
Billy Laws 70cffce5eb Windows: Avoid redeclaring strtoll as unimplemented
Included as part of the musl code we pull in. Fixes ARM64EC/WOW64 FEX.
2025-10-03 16:30:41 +02:00
Lioncache 0258e914e0 IR: Convert RegisterClassType to an enum class
Removes the last of the wrapper value types in favor of enum classes.
2025-10-03 03:54:56 -04:00
Ryan Houdek c35b81d2c5 Merge pull request #4925 from lioncash/regclass
IREmitter: Add register class helpers
2025-10-02 22:14:41 -07:00
Lioncache fc63dcedb6 IR: Migrate to new helpers
Reduces a bunch of noise related to the register classes and hoists them
out so that converting the classes over to enums should be fairly
straightforward.
2025-10-03 00:53:02 -04:00
Lioncache c879650c4b IR: Add GPR/FPR helpers
These will be utilized in follow up PRs
2025-10-02 23:53:50 -04:00
Ryan Houdek c4d019939b Merge pull request #4924 from lioncash/irdump 2025-10-02 10:36:12 -07:00
Lioncache 3841cc4aa5 IRDumper: Add remaining missing enum values
Now that the compiler can warn against these, we can fill the
remaining list in.
2025-10-02 12:18:36 -04:00
Lioncache 468f2f2d2d IRDumper: Convert remaining printers to lambda style invocations
This allows easily moving the default case out of the switch, making it
easier for compilers to warn about missing values in switches if any
enum members are added in the future but aren't added to the
formatters.
2025-10-02 12:18:28 -04:00
Lioncache be5db91c13 IRDumper: Add missing CheckTF BranchHint 2025-10-02 03:01:11 -04:00
Lioncache 6e7d0a520a IRDumper: Fix error message for FloatCompareOp
If ever printed this would give a misleading error that it was an
unrecognized OpSize type.
2025-10-02 02:57:29 -04:00
Lioncache e37996873a IRDumper: Add missing FenceType entry 2025-10-02 02:55:34 -04:00
Lioncache 25e851b570 IRDumper: Add missing SyscallFlags entry 2025-10-02 02:55:31 -04:00
Ryan Houdek d41b21a626 Merge pull request #4923 from lioncash/py5
IR: Remove TypeDefinition struct
2025-10-01 22:03:00 -07:00
Lioncache d65c54ba21 IR: Remove TypeDefinition struct
These aren't used at all.
2025-10-02 00:50:51 -04:00
Ryan Houdek a5d3bf0f48 Merge pull request #4922 from lioncash/py4
IR: Convert MemOffsetType to enum class
2025-10-01 21:41:54 -07:00
Lioncache f76e8c7185 IR: Convert MemOffsetType to enum class
Gets rid of another wrapper struct.
2025-10-02 00:24:42 -04:00
Ryan Houdek 96145e8e3f Merge pull request #4921 from lioncash/py3
IR: Convert rounding modes to enum class
2025-10-01 21:03:15 -07:00
Lioncache 530821fc4a IR: Convert rounding modes to enum class
Same behavior, but more compact
2025-10-01 23:37:57 -04:00
Ryan Houdek e4a4a529dd Merge pull request #4920 from lioncash/py2 2025-10-01 20:28:23 -07:00
Lioncache 863d1e0007 IR: Convert FenceType to an enum class
Same thing minus an extra struct lingering around.
2025-10-01 23:13:08 -04:00
Ryan Houdek c64d6f2b36 Merge pull request #4919 from lioncash/python
IR: Convert CondClassType over to enum class
2025-10-01 19:28:50 -07:00
Lioncache a798880ac8 IR: Convert CondClassType over to enum class
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.

This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.

Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
2025-10-01 10:45:00 -04:00
Ryan Houdek 9db9be9612 Merge pull request #4918 from lioncash/bitcast
General: Migrate to std::bit_cast
2025-09-30 19:32:51 -07:00
Lioncache 6a19c77184 General: Migrate to std::bit_cast
We already have a few cases where we already use bit_cast in the
emitter, so we may as well move all of our temporary helper instances
over as well.
2025-09-30 22:20:10 -04:00
Ryan Houdek c17b9dea30 Merge pull request #4917 from lioncash/collapse
StringConv: Merge integral-handling facilities together
2025-09-30 18:58:13 -07:00
Lioncache 7564b6c69a StringConv: Merge integral-handling facilities together
Same behavior, but we condense all of the integral handling
into one place.
2025-09-30 21:43:57 -04:00
Ryan Houdek 5858c5041f Merge pull request #4916 from lioncash/type
StringConv: Fix std::enable_if usage
2025-09-30 17:13:43 -07:00
LC 18d0dec416 Merge pull request #4911 from Sonicadvance1/evmd_linux
Linux: Implement support for extended volatile metadata
2025-09-30 20:01:26 -04:00
LC 752046108c Merge pull request #4910 from Sonicadvance1/fix_wow64_evmd
Wow64: Ensure extended volatile metadata is removed
2025-09-30 20:01:12 -04:00
Lioncache 773ddbf5c4 StringConv: Fix std::enable_if usage
This neglected to use the ::type qualifier, so this candidate was
always being considered in overload resolution.
2025-09-30 19:53:45 -04:00
Ryan Houdek 88e90a6f3a Merge pull request #4915 from lioncash/emitter
Arm64Emitter: Cull unnecessary includes
2025-09-30 10:50:04 -07:00
Lioncache 84325d6b5c Arm64Emitter: Cull unnecessary includes
Also fixes an indirect include.
2025-09-30 09:32:25 -04:00
Ryan Houdek 1b8e03e1de Merge pull request #4914 from lioncash/header
fextl/string: Correct <filesystem> include to <functional>
2025-09-29 21:26:39 -07:00
Lioncache 569d7297f0 fextl/string: Correct <filesystem> include to <functional>
This was unintentionally putting all the filesystem utilities into headers implicitly.
2025-09-29 23:54:08 -04:00
Ryan Houdek d9520c9498 Merge pull request #4913 from lioncash/help
JITClass: Mark functions as static where applicable
2025-09-29 20:53:09 -07:00
Lioncache 6b61093037 JITClass: Mark functions as static where applicable
These don't depend on any class state.
2025-09-29 23:42:43 -04:00
LC 74b09d5581 Merge pull request #4912 from Sonicadvance1/avx_32bit_overflow
FEXCore: Implement support for AVX gathers with overflow
2025-09-29 21:10:42 -04:00
Ryan Houdek 08da8f6f5c unittests/ASM: Adds gather overflow tests 2025-09-29 16:49:34 -07:00
Ryan Houdek 7f8fffbb24 FEXCore: Implement support for AVX gathers with overflow
Falls down the emulated path so we can zero extend the address
calculation. No way for SVE gathers to use 32-bit addressing.
2025-09-29 16:06:02 -07:00
Ryan Houdek b08e5f9821 Linux: Implement support for extended volatile metadata
Allows applications running entirely under Linux to use the same
extended volatile metadata as Windows.

For example `FEX_EXTENDEDVOLATILEMETADATA=iw4sp.exe\;0xe9da0-0xe9ec7`
this configuration works for both wow64 and Linux to disable the TSO
emulation on the memcpy routine in that game that consumes around 80% of
CPU time in TSO emulation.

Works with Linux native games as well of course.
2025-09-29 14:36:51 -07:00
Ryan Houdek 1758648f56 Common: Move ApplyFEXExtendedVolatileMetadata to Common
It is going to be used by Linux.

Also adds support for file offset handling, which is required on Linux
but not on win32.
2025-09-29 14:36:50 -07:00
Ryan Houdek 418ce0aed9 Wow64: Ensure extended volatile metadata is removed
Was missing this, matches the arm64ec side.
2025-09-29 14:13:44 -07:00
LC 4e926076c9 Merge pull request #4909 from zeyi2/main
Docs: Fix wrong link in `Readme_CN.md`
2025-09-29 10:30:38 -04:00
mtx 3f0aaeba68 Docs: Fix wrong link in Readme_CN.md 2025-09-29 16:53:31 +08:00
LC 6ecec77a74 Merge pull request #4908 from FEX-Emu/bylaws-patch-1
Change required Python version from 3.10 to 3.9
2025-09-28 13:03:45 -04:00
Billy Laws 79917dc9bb Change required Python version from 3.10 to 3.9
Debian bullseye is a supported build environment which only provides 3.9
2025-09-28 16:34:50 +02:00
Ryan Houdek b0b60fb761 Merge pull request #4907 from neobrain/feature_nix_update
Build: Update toolchain for WoA builds
2025-09-25 13:43:23 -07:00
Tony Wasserka 7463152f50 Build: Update toolchain for WoA builds 2025-09-25 22:17:44 +02:00
Ryan Houdek cbd8217c60 Merge pull request #4904 from bylaws/tesrgr
TestHarnessRunner: Don't attempt to build on MinGW
2025-09-24 13:16:47 -07:00
Ryan Houdek 4de1762c21 Merge pull request #4903 from bylaws/fee
vixl: Update submodule
2025-09-24 12:41:00 -07:00
Billy Laws 73962402df TestHarnessRunner: Don't attempt to build on MinGW 2025-09-24 20:28:33 +01:00
Billy Laws d90ccceec7 vixl: Update submodule 2025-09-24 20:26:28 +01:00
Ryan Houdek 0c5ef1ce1d Merge pull request #4902 from lioncash/fmt
Update fmt to 12.0.0
2025-09-23 12:13:34 -07:00
Ryan Houdek e5680e031d Merge pull request #4835 from pmatos/fix/classify
Fix quiet and signalling nan propagation
2025-09-23 11:17:37 -07:00
Paulo Matos eee0cffa57 instcountci: Fix quiet and signalling nan propagation 2025-09-23 13:48:45 +02:00
Paulo Matos c480b0ba41 asm_tests: Fix quiet and signalling nan propagation 2025-09-23 13:48:45 +02:00
Paulo Matos e7a47a647c Fix quiet and signalling nan propagation
This adds a new mode X87StrictReducedPrecision.
The strict reduced precision is like the reduced precision but adds extra checks,
like the currently implemented nan and snan propagations.

Fix for __builtin_issignaling() test of SPEC2017 classify test.
2025-09-23 13:48:45 +02:00
Paulo Matos 500c82bbf8 Do not format NASM include files 2025-09-23 13:48:45 +02:00
Lioncache 008528af91 Update fmt to 12.0.0
Moves us over to the next major version release.
Changelog here: https://github.com/fmtlib/fmt/releases/tag/12.0.0
2025-09-23 09:50:11 +02:00
Tony Wasserka 13199ad203 Merge pull request #4898 from Sonicadvance1/a_snake_a_snake
CMake: Update our python minspec to 3.10
2025-09-23 09:29:33 +02:00
LC 533dd1a54b Merge pull request #4897 from Sonicadvance1/fix_telem
OpcodeDispatcher: Fixes telemetry on legacy segment read
2025-09-23 08:11:10 +02:00
Ryan Houdek b6f2e10416 CMake: Update our python minspec to 3.10 2025-09-22 12:31:29 -07:00
Ryan Houdek c58db36a7c OpcodeDispatcher: Fixes telemetry on legacy segment read 2025-09-22 11:57:34 -07:00
727 changed files with 44164 additions and 37489 deletions

No files matched your search

+3
View File
@@ -7,3 +7,6 @@ FEXCore/Source/Interface/Core/X86Tables/*
# Inline headers with list-like content that can't be processed individually
Source/Tools/LinuxEmulation/LinuxSyscalls/x*/SyscallsNames.inl
Source/Tools/LinuxEmulation/LinuxSyscalls/x*/Ioctl/*.inl
# Include files in unittests
unittests/*ASM/Includes/*.inc
+79
View File
@@ -0,0 +1,79 @@
name: steamrt4 build
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
jobs:
steamrt4_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, distrobox]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name : submodule checkout
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build
cmake -E make_directory ${{runner.workspace}}/build
# Setup everything required.
- name : distrobox setup
run: |
distrobox create -Y -i registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306 steamrt4 || true
distrobox upgrade steamrt4
distrobox enter --name steamrt4 -- sudo apt-get install -y \
git cmake ninja-build ccache \
lld clang \
libclang-dev llvm-dev \
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross \
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- name: Create Build Environment
run: distrobox enter --name steamrt4 -- cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
shell: bash
working-directory: ${{runner.workspace}}/build
run: distrobox enter --name steamrt4 -- cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE
- name: install
working-directory: ${{runner.workspace}}/build
shell: bash
env:
DESTDIR: ${{runner.workspace}}/install
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE -t install
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
overwrite: true
name: steamrt4_steampipe_depot
path: ${{runner.workspace}}/install/*
retention-days: 1
compression-level: 9
+15 -3
View File
@@ -33,6 +33,7 @@ option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling ca
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use for the FEXCore profiler (gpuvis, tracy)")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
option(BUILD_STEAM_SUPPORT "Builds FEX for integration into Steam" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
@@ -64,6 +65,10 @@ if (NOT MINGW_BUILD)
endif()
endif()
if (BUILD_STEAM_SUPPORT)
add_definitions(-DFEX_STEAM_SUPPORT=1)
endif()
if (ENABLE_FEXCORE_PROFILER)
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
@@ -319,7 +324,7 @@ if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
find_package(Python 3.9 REQUIRED COMPONENTS Interpreter)
set(BUILD_SHARED_LIBS OFF)
@@ -476,13 +481,16 @@ add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
if (_M_ARM_64 AND NOT MINGW_BUILD)
if (_M_ARM_64 AND NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
# Binfmt_misc files must be installed prior to Source/ installs
add_subdirectory(Data/binfmts/)
endif()
add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
if (NOT BUILD_STEAM_SUPPORT)
add_subdirectory(Data/AppConfig/)
endif()
# Install the ThunksDB file
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS ${CMAKE_CURRENT_SOURCE_DIR}/Data/*.json)
@@ -582,6 +590,10 @@ if (BUILD_THUNKS)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
if (BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
set(FEX_VERSION_MAJOR "0")
set(FEX_VERSION_MINOR "0")
set(FEX_VERSION_PATCH "0")
+21 -13
View File
@@ -38,9 +38,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded adr(ARMEmitter::Register rd, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
if (IsADRRange(Imm)) [[likely]] {
if (IsADRRange(Imm)) {
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
@@ -73,9 +71,8 @@ public:
[[nodiscard]] BranchEncodeSucceeded adrp(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) [[likely]] {
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) {
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
@@ -103,16 +100,22 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>());
const auto SLocation = reinterpret_cast<int64_t>(Label->Location);
const auto ULocation = std::bit_cast<uint64_t>(SLocation);
const int64_t Imm = SLocation - (GetCursorAddress<int64_t>());
const auto UImm = std::bit_cast<uint64_t>(Imm);
if (IsADRRange(Imm)) {
// If the range is in ADR range then we can just use ADR.
return adr(rd, Label);
} else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
}
if (IsADRPRange(Imm)) {
const int64_t ADRPImm = (SLocation & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
// If the range is in the ADRP range then we can use ADRP.
bool NeedsOffset = !IsADRPAligned(reinterpret_cast<uint64_t>(Label->Location));
uint64_t AlignedOffset = reinterpret_cast<uint64_t>(Label->Location) & 0xFFFULL;
const bool NeedsOffset = !IsADRPAligned(ULocation);
const uint64_t AlignedOffset = ULocation & 0xFFFULL;
// First emit ADRP
adrp(rd, ADRPImm >> 12);
@@ -125,14 +128,19 @@ public:
return BranchEncodeSucceeded::Success;
}
// Can't encode.
return BranchEncodeSucceeded::Failure;
// Stinky path, we need to load the address as a sequence of movz+movk+movk
movz(ARMEmitter::Size::i64Bit, rd, (UImm >> 32) & 0xFFFF, 32);
movk(ARMEmitter::Size::i64Bit, rd, (UImm >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, rd, UImm & 0xFFFF);
return BranchEncodeSucceeded::Success;
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::LONG_ADDRESS_GEN});
// Emit a register index and a nop. These will be backpatched.
// Emit a register index and two nops. These will be backpatched.
dc32(rd.Idx());
nop();
nop();
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
+8 -9
View File
@@ -22,7 +22,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded b(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -55,7 +55,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded bc(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -116,7 +116,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded b(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -151,7 +151,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded bl(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
@@ -189,7 +189,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded cbz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -227,7 +227,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -265,7 +265,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded tbz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -301,9 +301,8 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded tbnz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) [[likely]] {
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
+31 -21
View File
@@ -586,6 +586,10 @@ concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegi
template<typename T>
concept IsQOrDRegister = std::is_same_v<T, QRegister> || std::is_same_v<T, DRegister>;
template<typename T>
concept IsLabel = std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>;
enum class BranchEncodeSucceeded {
Success,
Failure,
@@ -658,7 +662,7 @@ public:
case ForwardLabel::InstType::ADR: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!IsADRRange(Imm)) [[unlikely]] {
if (!IsADRRange(Imm)) {
// Can't bind.
return false;
}
@@ -674,7 +678,7 @@ public:
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(IsADRPRange(Imm) && IsADRPAligned(Imm))) [[unlikely]] {
if (!(IsADRPRange(Imm) && IsADRPAligned(Imm))) {
// Can't bind.
return false;
}
@@ -691,7 +695,7 @@ public:
case ForwardLabel::InstType::B: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0))) [[unlikely]] {
if (!(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0))) {
// Can't bind.
return false;
}
@@ -707,7 +711,7 @@ public:
case ForwardLabel::InstType::TEST_BRANCH: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0))) [[unlikely]] {
if (!(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0))) {
// Can't bind.
return false;
}
@@ -724,7 +728,7 @@ public:
case ForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0))) [[unlikely]] {
if (!(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0))) {
// Can't bind.
return false;
}
@@ -737,38 +741,44 @@ public:
break;
}
case ForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
const auto* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
const auto ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
const auto ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
const auto ImmInstThree = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[2]);
const auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
const auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
if (IsADRRange(ImmInstThree)) {
// If within ADR range from the third instruction, then we can emit NOP+NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
} else if (IsADRPRange(ImmInstOne)) {
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstThree) & 0x7FFF);
} else if (IsADRPRange(ImmInstTwo)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
// We can emit nop + nop + adrp
nop();
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstThree >> 12) & 0x7FFF);
} else {
// Not aligned, need nop + adrp + add
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
} else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstTwo & 0xFFF);
}
} else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
// Stinky path, we need to emit a movz+movk+movk sequence.
movz(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 32) & 0x7FFF, 32);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne) & 0xFFFF);
}
SetCursorOffset(OriginalOffset);
+3 -3
View File
@@ -5125,7 +5125,7 @@ private:
requires (std::is_same_v<T, float> || std::is_same_v<T, double>)
[[nodiscard]]
static bool IsValidFPValueForImm8(T value) {
const uint64_t bits = FEXCore::BitCast<FloatToEquivalentUInt<T>>(value);
const uint64_t bits = std::bit_cast<FloatToEquivalentUInt<T>>(value);
const uint64_t datasize_idx = FEXCore::ilog2(sizeof(T)) - 1;
static constexpr std::array mantissa_masks {
@@ -5171,7 +5171,7 @@ protected:
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value), "Value ({}) cannot be encoded into an 8-bit immediate", value);
#endif
const auto bits = FEXCore::BitCast<uint32_t>(value);
const auto bits = std::bit_cast<uint32_t>(value);
const auto sign = (bits & 0x80000000) >> 24;
const auto expb2 = (bits & 0x20000000) >> 23;
const auto b5_to_0 = (bits >> 19) & 0x3F;
@@ -5184,7 +5184,7 @@ protected:
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value), "Value ({}) cannot be encoded into an 8-bit immediate", value);
#endif
const auto bits = FEXCore::BitCast<uint64_t>(value);
const auto bits = std::bit_cast<uint64_t>(value);
const auto sign = (bits & 0x80000000'00000000) >> 56;
const auto expb2 = (bits & 0x20000000'00000000) >> 55;
const auto b5_to_0 = (bits >> 48) & 0x3F;
+2 -2
View File
@@ -2,8 +2,8 @@
let
toolchain = pkgs.fetchzip {
url = "https://github.com/bylaws/llvm-mingw/releases/download/20250305/llvm-mingw-20250305-ucrt-ubuntu-20.04-aarch64.tar.xz";
sha256 = "sha256-cA03/ab9O61eO9+S2JzIXD4V0HzTXK5/AYyxW2d73Po=";
url = "https://github.com/bylaws/llvm-mingw/releases/download/20250920/llvm-mingw-20250920-ucrt-ubuntu-22.04-aarch64.tar.xz";
sha256 = "sha256-LaojKjC8KzY+soW5u6eoDoXE3qtYk9Ejr7M3enTqRAE=";
};
cmakeToolchainFile = pkgs.substitute {
+1 -1
+1 -1
+1 -1
+5 -3
View File
@@ -74,9 +74,11 @@ add_compile_options($<$<COMPILE_LANGUAGE:CXX>:-fno-strict-aliasing> $<$<COMPILE_
add_subdirectory(Source/)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
if (NOT BUILD_STEAM_SUPPORT)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
endif()
if (BUILD_TESTING)
add_subdirectory(unittests/)
+9
View File
@@ -200,6 +200,15 @@ def print_man_environment_tail():
],
"''", True)
print_man_env_option(
"APP_CACHE_LOCATION",
[
"Allows the user to override where FEX stores and loads cache files",
"By default FEX will look in $XDG_CACHE_HOME/fex-emu/ or $HOME/.cache/fex-emu/",
"This will override the full path, trailing forward-slash is expected to exist",
],
"''", True)
def print_man_header():
header ='''.Dd {0}
.Dt FEX
+53 -59
View File
@@ -58,10 +58,10 @@ class OpDefinition:
JITDispatch: bool
JITDispatchOverride: str
TiedSource: int
Inline: list
Arguments: list
EmitValidation: list
Desc: list
Inline: list[str]
Arguments: list[OpArgument]
EmitValidation: list[str]
Desc: list[str]
def __init__(self):
self.Name = None
@@ -92,19 +92,14 @@ class OpDefinition:
attrs = vars(self)
print(", ".join("%s: %s" % item for item in attrs.items()))
IRTypesToCXX = {}
CXXTypeToIR = {}
IROps = []
IRTypesToCXX: dict[str, IRType] = {}
CXXTypeToIR: dict[str, IRType] = {}
IROps: list[OpDefinition] = []
IROpNameMap = {}
IROpNameSet: set[str] = set()
def is_ssa_type(type):
if (type == "SSA" or
type == "GPR" or
type == "GPRPair" or
type == "FPR"):
return True
return False
def is_ssa_type(op_type: str):
return op_type in {"SSA", "GPR", "GPRPair", "FPR"}
def parse_irtypes(irtypes):
for op_key, op_val in irtypes.items():
@@ -219,11 +214,8 @@ def parse_ops(ops):
OpArg.DefaultInitializer = DefaultInit[1][:-1]
# If SSA type then we can generate validation for this op
if (OpArg.IsSSA and
(OpArg.Type == "GPR" or
OpArg.Type == "GPRPair" or
OpArg.Type == "FPR")):
OpDef.EmitValidation.append(f"GetOpRegClass({ArgName}) == InvalidClass || WalkFindRegClass({ArgName}) == {OpArg.Type}Class")
if OpArg.IsSSA and OpArg.Type in {"GPR", "GPRPair", "FPR"}:
OpDef.EmitValidation.append(f"GetOpRegClass({ArgName}) == RegClass::Invalid || WalkFindRegClass({ArgName}) == RegClass::{OpArg.Type}")
OpArg.Name = ArgName
OpArg.NameWithPrefix = NameWithPrefix
@@ -296,22 +288,29 @@ def parse_ops(ops):
#OpDef.print()
# Error on duplicate op
if OpDef.Name in IROpNameMap:
if OpDef.Name in IROpNameSet:
ExitError("Duplicate Op defined! {}".format(OpDef.Name))
IROps.append(OpDef)
IROpNameMap[OpDef.Name] = 1
IROpNameSet.add(OpDef.Name)
# Print out enum values
def print_enums():
def print_enums(enums):
output_file.write("#ifdef IROP_ENUM\n")
output_file.write("enum IROps : uint16_t {\n")
for op in IROps:
output_file.write("\tOP_{},\n" .format(op.Name.upper()))
output_file.write("};\n")
for name, members in enums.items():
output_file.write(f"enum {name} {{\n")
for member in members:
if member:
output_file.write(f"\t{member}\n")
else:
output_file.write("\n")
output_file.write("};\n\n")
output_file.write("#undef IROP_ENUM\n")
output_file.write("#endif\n\n")
@@ -408,7 +407,7 @@ def print_ir_sizes():
[[nodiscard, gnu::const]] std::string_view const& GetName(IROps Op);
[[nodiscard, gnu::const]] uint8_t GetArgs(IROps Op);
[[nodiscard, gnu::const]] uint8_t GetRAArgs(IROps Op);
[[nodiscard, gnu::const]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);
[[nodiscard, gnu::const]] FEXCore::IR::RegClass GetRegClass(IROps Op);
[[nodiscard, gnu::const]] bool HasSideEffects(IROps Op);
[[nodiscard, gnu::const]] bool ImplicitFlagClobber(IROps Op);
[[nodiscard, gnu::const]] bool GetHasDest(IROps Op);
@@ -422,30 +421,29 @@ def print_ir_sizes():
def print_ir_reg_classes():
output_file.write("#ifdef IROP_REG_CLASSES_IMPL\n")
output_file.write("constexpr std::array<FEXCore::IR::RegisterClassType, IROps::OP_LAST + 1> IRRegClasses = {\n")
output_file.write("constexpr std::array<FEXCore::IR::RegClass, IROps::OP_LAST + 1> IRRegClasses = {\n")
for op in IROps:
if op.Name == "Last":
output_file.write("\tFEXCore::IR::InvalidClass,\n")
output_file.write("\tRegClass::Invalid,\n")
else:
Class = "Invalid"
if op.HasDest and op.DestType == None:
if op.HasDest and op.DestType is None:
ExitError("IR op {} has destination with no destination class".format(op.Name))
if op.HasDest and op.DestType == "SSA": # Special case SSA type
output_file.write("\tFEXCore::IR::ComplexClass,\n")
output_file.write("\tRegClass::Complex,\n")
elif op.HasDest:
output_file.write("\tFEXCore::IR::{}Class,\n".format(op.DestType))
output_file.write("\tRegClass::{},\n".format(op.DestType))
else:
# No destination so it has an invalid destination class
output_file.write("\tFEXCore::IR::InvalidClass, // No destination\n")
output_file.write("\tRegClass::Invalid, // No destination\n")
output_file.write("};\n\n")
output_file.write("// Make sure our array maps directly to the IROps enum\n")
output_file.write("static_assert(IRRegClasses[IROps::OP_LAST] == FEXCore::IR::InvalidClass);\n\n")
output_file.write("static_assert(IRRegClasses[IROps::OP_LAST] == RegClass::Invalid);\n\n")
output_file.write("FEXCore::IR::RegisterClassType GetRegClass(IROps Op) { return IRRegClasses[Op]; }\n\n")
output_file.write("FEXCore::IR::RegClass GetRegClass(IROps Op) { return IRRegClasses[Op]; }\n\n")
output_file.write("#undef IROP_REG_CLASSES_IMPL\n")
output_file.write("#endif\n\n")
@@ -568,9 +566,7 @@ def print_ir_arg_printer():
SSAArgNum = 0
FirstArg = True
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
for arg in op.Arguments:
# No point printing temporaries that we can't recover
if arg.Temporary:
continue
@@ -671,7 +667,7 @@ def print_ir_allocator_helpers():
output_file.write("\t\treturn HeaderOp->Op;\n")
output_file.write("\t}\n\n")
output_file.write("\tFEXCore::IR::RegisterClassType GetOpRegClass(const OrderedNode *Op) const {\n")
output_file.write("\tFEXCore::IR::RegClass GetOpRegClass(const OrderedNode *Op) const {\n")
output_file.write("\t\treturn GetRegClass(GetOpType(Op));\n")
output_file.write("\t}\n\n")
@@ -685,22 +681,21 @@ def print_ir_allocator_helpers():
output_file.write("\tIRPair<IROp_{}> _{}(" .format(op.Name, op.Name))
# Output SSA args first
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
for i, arg in enumerate(op.Arguments):
LastArg = i == len(op.Arguments) - 1
if arg.Temporary:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
elif arg.IsSSA:
# SSA value
output_file.write("OrderedNodeWrapper {}".format(arg.Name))
else:
# User defined op that is stored
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
if arg.DefaultInitializer != None:
if arg.DefaultInitializer:
output_file.write(" = {}".format(arg.DefaultInitializer))
if not LastArg:
@@ -758,20 +753,19 @@ def print_ir_allocator_helpers():
if op.SSAArgNum:
output_file.write("\tIRPair<IROp_{}> _{}(" .format(op.Name, op.Name))
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
for i, arg in enumerate(op.Arguments):
LastArg = i == len(op.Arguments) - 1
if arg.Temporary:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
elif arg.IsSSA:
output_file.write("OrderedNode *{}".format(arg.Name))
else:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
if arg.DefaultInitializer != None:
if arg.DefaultInitializer:
output_file.write(" = {}".format(arg.DefaultInitializer))
if not LastArg:
@@ -812,16 +806,15 @@ def print_ir_allocator_helpers():
print_validation(op)
output_file.write(f"\t\treturn _{op.Name}(")
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
for i, arg in enumerate(op.Arguments):
LastArg = i == len(op.Arguments) - 1
output_file.write(arg.Name)
if arg.IsSSA:
output_file.write("->Wrapped(ListDataBegin)")
if not LastArg:
output_file.write(", ")
output_file.write(");\n");
output_file.write("\t}\n\n");
output_file.write(");\n")
output_file.write("\t}\n\n")
output_file.write("#undef IROP_ALLOCATE_HELPERS\n")
output_file.write("#endif\n")
@@ -852,8 +845,8 @@ def print_ir_dispatcher_dispatch():
output_dispatch_file.write("#endif\n")
if (len(sys.argv) < 4):
ExitError()
if len(sys.argv) < 4:
ExitError("Insufficient parameters passed to script")
output_filename = sys.argv[2]
output_dispatcher_filename = sys.argv[3]
@@ -865,6 +858,7 @@ json_file.close()
json_object = json.loads(json_text)
json_object = {k.upper(): v for k, v in json_object.items()}
enums = json_object["ENUMS"]
ops = json_object["OPS"]
irtypes = json_object["IRTYPES"]
defines = json_object["DEFINES"]
@@ -874,7 +868,7 @@ parse_ops(ops)
output_file = open(output_filename, "w")
print_enums()
print_enums(enums)
print_ir_structs(defines)
print_ir_sizes()
print_ir_reg_classes()
+5 -4
View File
@@ -31,7 +31,6 @@ set (SRCS
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
@@ -202,8 +201,10 @@ add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_MAN_NAME}"
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
if (NOT BUILD_STEAM_SUPPORT)
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
endif()
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
@@ -290,7 +291,7 @@ AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
if (NOT MINGW_BUILD)
if (NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
install(TARGETS ${PROJECT_NAME}_shared
LIBRARY
DESTINATION ${CMAKE_INSTALL_LIBDIR}
+3 -2
View File
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/memory.h>
@@ -12,7 +13,7 @@ namespace FEXCore {
// Buffered JIT symbol tracking.
struct JITSymbolBuffer {
// Maximum buffer size to ensure we are a page in size.
constexpr static size_t BUFFER_SIZE = 4096 - (8 * 2);
constexpr static size_t BUFFER_SIZE = FEXCore::Utils::FEX_PAGE_SIZE - (8 * 2);
// Maximum distance until the end of the buffer to do a write.
constexpr static size_t NEEDS_WRITE_DISTANCE = BUFFER_SIZE - 64;
// Maximum time threshhold to wait before a buffer write occurs.
@@ -27,7 +28,7 @@ struct JITSymbolBuffer {
size_t Offset {};
char Buffer[BUFFER_SIZE] {};
};
static_assert(sizeof(JITSymbolBuffer) == 4096, "Ensure this is one page in size");
static_assert(sizeof(JITSymbolBuffer) == FEXCore::Utils::FEX_PAGE_SIZE, "Ensure this is one page in size");
class JITSymbols final {
public:
+8 -8
View File
@@ -4,9 +4,9 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/BitUtils.h>
#include "cephes_128bit.h"
#include <bit>
#include <cmath>
#include <cstring>
#include <stdint.h>
@@ -501,12 +501,12 @@ struct FEX_PACKED X80SoftFloat {
float ToF32(softfloat_state* state) const {
const float32_t Result = extF80_to_f32(state, *this);
return FEXCore::BitCast<float>(Result);
return std::bit_cast<float>(Result);
}
double ToF64(softfloat_state* state) const {
const float64_t Result = extF80_to_f64(state, *this);
return FEXCore::BitCast<double>(Result);
return std::bit_cast<double>(Result);
}
FEXCore::VectorRegType ToVector() const {
@@ -518,7 +518,7 @@ struct FEX_PACKED X80SoftFloat {
BIGFLOAT ToFMax(softfloat_state* state) const {
#if BIGFLOATSIZE == 16
const float128_t Result = extF80_to_f128(state, *this);
return FEXCore::BitCast<BIGFLOAT>(Result);
return std::bit_cast<BIGFLOAT>(Result);
#else
BIGFLOAT result {};
memcpy(&result, this, sizeof(result));
@@ -577,18 +577,18 @@ struct FEX_PACKED X80SoftFloat {
}
X80SoftFloat(softfloat_state* state, const float rhs) {
*this = f32_to_extF80(state, FEXCore::BitCast<float32_t>(rhs));
*this = f32_to_extF80(state, std::bit_cast<float32_t>(rhs));
}
X80SoftFloat(softfloat_state* state, const double rhs) {
*this = f64_to_extF80(state, FEXCore::BitCast<float64_t>(rhs));
*this = f64_to_extF80(state, std::bit_cast<float64_t>(rhs));
}
X80SoftFloat(softfloat_state* state, BIGFLOAT rhs) {
#if BIGFLOATSIZE == 16
*this = f128_to_extF80(state, FEXCore::BitCast<float128_t>(rhs));
*this = f128_to_extF80(state, std::bit_cast<float128_t>(rhs));
#else
*this = FEXCore::BitCast<long double>(rhs);
*this = std::bit_cast<long double>(rhs);
#endif
}
+11 -47
View File
@@ -2,59 +2,23 @@
#pragma once
#include <FEXCore/fextl/string.h>
#include <cstdint>
#include <concepts>
#include <string_view>
#include <optional>
namespace FEXCore::StrConv {
inline bool Conv(std::string_view Value, bool* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
template<std::integral T>
bool Conv(std::string_view Value, T* Result) {
if constexpr (std::is_signed_v<T>) {
*Result = static_cast<T>(std::strtoll(Value.data(), nullptr, 0));
} else {
*Result = static_cast<T>(std::strtoull(Value.data(), nullptr, 0));
}
return true;
}
inline bool Conv(std::string_view Value, uint8_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
inline bool Conv(std::string_view Value, int8_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
inline bool Conv(std::string_view Value, uint16_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
inline bool Conv(std::string_view Value, int16_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
inline bool Conv(std::string_view Value, uint32_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
inline bool Conv(std::string_view Value, int32_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
inline bool Conv(std::string_view Value, uint64_t* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
inline bool Conv(std::string_view Value, int64_t* Result) {
*Result = std::strtoll(Value.data(), nullptr, 0);
return true;
}
template<typename T, typename = std::enable_if<std::is_enum<T>::value, T>>
inline bool Conv(std::string_view Value, T* Result) {
*Result = static_cast<T>(std::stoull(Value.data(), nullptr, 0));
template<typename T, typename = std::enable_if_t<std::is_enum_v<T>, T>>
bool Conv(std::string_view Value, T* Result) {
*Result = static_cast<T>(std::strtoull(Value.data(), nullptr, 0));
return true;
}
+19 -14
View File
@@ -30,14 +30,14 @@ class Context;
}
namespace FEXCore::Config {
namespace DefaultValues {
namespace detail {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) const P(type) P(enum) = P(default);
#define OPT_STR(group, enum, json, default) const std::string_view P(enum) = P(default);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#define OPT_STRENUM(group, enum, json, default) const uint64_t P(enum) = FEXCore::ToUnderlying(P(default));
#include <FEXCore/Config/ConfigValues.inl>
} // namespace DefaultValues
} // namespace detail
enum Paths {
PATH_DATA_DIR_LOCAL = 0,
@@ -134,7 +134,7 @@ public:
void Load();
template<typename T>
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<DefaultValues::Type::StringArrayType, T>)
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<StringArrayType, T>)
std::optional<T> GetConv(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
@@ -142,7 +142,7 @@ public:
}
const auto& Value = it->second;
LOGMAN_THROW_A_FMT(!std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(!std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
if (std::holds_alternative<T>(Value)) [[likely]] {
return std::get<T>(Value);
@@ -165,7 +165,7 @@ public:
private:
void MergeConfigMap(const LayerOptions& Options);
void MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value);
void MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value);
};
void MetaLayer::Load() {
@@ -181,7 +181,7 @@ void MetaLayer::Load() {
}
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value) {
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
@@ -193,7 +193,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](const DefaultValues::Type::StringArrayType& Value) {
const auto AddToMap = [&LookupMap](const StringArrayType& Value) {
for (const auto& EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == fextl::string::npos) {
@@ -209,7 +209,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
}
};
AddToMap(std::get<DefaultValues::Type::StringArrayType>(MetaEnvironment->second));
AddToMap(std::get<StringArrayType>(MetaEnvironment->second));
AddToMap(Value);
// Now with the two layers merged in the map
@@ -225,8 +225,8 @@ void MetaLayer::MergeConfigMap(const LayerOptions& Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto& it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV || it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<DefaultValues::Type::StringArrayType>(it.second));
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<StringArrayType>(it.second));
} else {
OptionMap.insert_or_assign(it.first, it.second);
}
@@ -423,7 +423,7 @@ bool Exists(ConfigOption Option) {
return Meta->OptionExists(Option);
}
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
std::optional<StringArrayType*> All(ConfigOption Option) {
return Meta->All(Option);
}
@@ -436,6 +436,12 @@ std::optional<T> GetConv(ConfigOption Option) {
return Meta->GetConv<T>(Option);
}
template std::optional<bool> GetConv(ConfigOption Option);
template std::optional<uint8_t> GetConv(ConfigOption Option);
template std::optional<int32_t> GetConv(ConfigOption Option);
template std::optional<uint32_t> GetConv(ConfigOption Option);
template std::optional<uint64_t> GetConv(ConfigOption Option);
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
@@ -491,13 +497,12 @@ template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t De
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List) {
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<DefaultValues::Type::StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option,
DefaultValues::Type::StringArrayType* List);
template void Value<StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
} // namespace FEXCore::Config
+67 -22
View File
@@ -16,6 +16,13 @@
"Maximum number of instruction to store in a block"
]
},
"EnableCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable the code caching subsystem"
]
},
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
@@ -59,7 +66,9 @@
"ENABLEWFXT": "enablewfxt",
"DISABLEWFXT": "disablewfxt",
"ENABLE3DNOW": "enable3dnow",
"DISABLE3DNOW": "disable3dnow"
"DISABLE3DNOW": "disable3dnow",
"ENABLESSE4A": "enablesse4a",
"DISABLESSE4A": "disablesse4a"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -82,7 +91,8 @@
"\t{enable,disable}svebitperm: Will force enable or disable svebitperm even if the host doesn't support it",
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it",
"\t{enable,disable}3dnow: Will force enable or disable 3DNow even if the host doesn't support it"
"\t{enable,disable}3dnow: Will force enable or disable 3DNow! even if the host doesn't support it",
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -91,6 +101,13 @@
"Desc": [
"Scales the cycle counter on systems that have low frequencies."
]
},
"CPUFeatureRegisters": {
"Type": "str",
"Default": "",
"Desc": [
"Allows overriding cpu feature flags for manual testing"
]
}
},
"Emulation": {
@@ -158,6 +175,44 @@
"Desc": [
"Allows the user to pass additional arguments to the application"
]
},
"DisableL2Cache": {
"Type": "bool",
"Default": "false",
"Desc": [
"Disables FEXCore's JIT L2 cache lookup. Saving memory.",
"Can potentially introduce more stutters."
]
},
"DynamicL1Cache": {
"Type": "bool",
"Default": "false",
"Desc": [
"Switches FEXCore's JIT L1 cache to be dynamically sized. Saving memory.",
"Can potentially introduce more stutters."
]
},
"DynamicL1CacheIncreaseCountHeuristic": {
"Type": "uint64",
"Default": "250",
"Desc": [
"Threshold of lookups per second that the L1 dynamic cache should increase its size.",
"Lower numbers means more aggressive scaling upward to the maximum size.",
"Higher numbers means more conservative scaling, using less memory.",
"Can potentially introduce stutters, more likely the higher the number.",
"Don't have this number smaller than the decrease count!"
]
},
"DynamicL1CacheDecreaseCountHeuristic": {
"Type": "uint64",
"Default": "50",
"Desc": [
"Threshold of lookups per second that the L1 dynamic cache should decrease its size.",
"The higher the number, the more aggressively it reduces the L1 cache size.",
"Lower numbers means more conservative memory savings.",
"Can potentially introduce more stutters, more likely the higher the number.",
"Don't have this number larger than the increase count!"
]
}
},
"Debug": {
@@ -328,11 +383,11 @@
"Requires a supported version of Mangohud to see the results"
]
},
"TraceProfiler": {
"EnableGpuvisProfiling": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enables FEX's trace profiler. Using gpuvis or tracy"
"Enables profiling when FEX was built with the gpuvis profiler backend."
]
}
},
@@ -388,12 +443,19 @@
"This is required to ensure a split-lock doesn't tear inside the process"
]
},
"KernelUnalignedAtomicBackpatching": {
"Type": "bool",
"Default": "true",
"Desc": [
"When the kernel unaligned atomic handler is enabled, use backpatching to reduce kernel context switches."
]
},
"VolatileMetadata": {
"Type": "bool",
"Default": "true",
"Desc": [
"Use volatile metadata in PE files to inform TSO instructions when available.",
"When metadata is unavailable falls back to the currently enabled TSO options."
"When metadata is unavailable falls back to the currently enabled TSO options."
]
},
"X87ReducedPrecision": {
@@ -403,23 +465,6 @@
"Emulates X87 floating point using 64-bit precision. This reduces emulation accuracy and may result in rendering bugs."
]
},
"ABILocalFlags": {
"Type": "bool",
"Default": "false",
"Desc": [
"When enabled enables an optimization around flags.",
"Assumes flags are not used across cals.",
"Hand-written assembly can violate this assumption."
]
},
"ParanoidTSO": {
"Type": "bool",
"Default": "false",
"Desc": [
"Makes TSO operations even more strict.",
"Forces vector loadstores to also become atomic."
]
},
"StallProcess": {
"Type": "bool",
"Default": "false",
+44 -17
View File
@@ -4,7 +4,6 @@
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -29,6 +28,7 @@
namespace FEXCore {
class SignalDelegator;
class ThunkHandler;
struct LookupCacheWriteLockToken;
namespace Core {
struct DebugData;
@@ -61,7 +61,7 @@ struct CustomIRResult {
, Data(Data) {}
};
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class CodeCache : public AbstractCodeCache {
@@ -72,12 +72,34 @@ public:
ContextImpl& CTX;
bool IsGeneratingCache = false;
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
void LoadData(Core::InternalThreadState&, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
void InitiateCacheGeneration() override {
IsGeneratingCache = true;
}
/**
* Applies a set of FEX relocations to the given code section.
*
* FEX relocations describe runtime-dependencies of FEX-generated code.
* When loading a code cache, they are used to move cached code to the
* dynamically chosen base address of the guest binary.
*
* Conversely, relocations are applied in reverse when writing code caches
* to ensure consistency across generation runs.
*
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
@@ -154,10 +176,20 @@ public:
return CodeCache;
}
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter> Writer) override {
CodeMapWriter = std::move(Writer);
}
void FlushAndCloseCodeMap() override {
if (CodeMapWriter) {
CodeMapWriter.reset();
}
}
void OnCodeBufferAllocated(const std::shared_ptr<CPU::CodeBuffer>&) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start,
uint64_t Length) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -197,7 +229,6 @@ public:
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(VectorTSOEnabled, VECTORTSOENABLED);
FEX_CONFIG_OPT(MemcpySetTSOEnabled, MEMCPYSETTSOENABLED);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(SMCChecks, SMCCHECKS);
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
@@ -205,7 +236,6 @@ public:
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(GDBSymbols, GDBSYMBOLS);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
@@ -226,14 +256,12 @@ public:
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CodeCache CodeCache;
fextl::unique_ptr<CodeMapWriter> CodeMapWriter;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl(const FEXCore::HostFeatures& Features);
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
// This is used as a replacement for the SMC writes in the mono callsite backpatcher that avoids atomic operations
@@ -268,9 +296,9 @@ public:
FEXCore::JITSymbols Symbols;
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator;
FEXCore::Utils::PooledAllocatorVirtual CPUBackendAllocator;
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator {"FEXMem_OpDispatcher"};
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator {"FEXMem_Frontend"};
FEXCore::Utils::PooledAllocatorVirtualWithGuard CPUBackendAllocator {"FEXMem_CPUBackend"};
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const {
@@ -311,10 +339,6 @@ protected:
AtomicTSOEmulationEnabled = false;
VectorAtomicTSOEmulationEnabled = false;
MemcpyAtomicTSOEmulationEnabled = false;
} else if (Config.ParanoidTSO) {
AtomicTSOEmulationEnabled = true;
VectorAtomicTSOEmulationEnabled = true;
MemcpyAtomicTSOEmulationEnabled = true;
} else {
AtomicTSOEmulationEnabled = Config.TSOEnabled;
VectorAtomicTSOEmulationEnabled = Config.TSOEnabled && Config.VectorTSOEnabled;
@@ -353,5 +377,8 @@ private:
bool MonoDetected = false;
std::atomic<uint64_t> MonoBackpatcherBlock;
std::mutex CodeBufferListLock;
fextl::vector<std::weak_ptr<CPU::CodeBuffer>> CodeBufferList;
};
} // namespace FEXCore::Context
+13 -11
View File
@@ -7,7 +7,7 @@
namespace FEXCore::IR {
Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage) {
Ref LoadEffectiveAddress(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage) {
Ref Tmp = A.Base;
if (A.Offset) {
@@ -51,8 +51,8 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
return Tmp ?: IREmit->Constant(0);
}
AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO, bool Vector,
IR::OpSize AccessSize) {
AddressMode SelectAddressMode(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO,
bool Vector, IR::OpSize AccessSize) {
const auto Is32Bit = GPRSize == OpSize::i32Bit;
const auto GPRSizeMatchesAddrSize = A.AddrSize == GPRSize;
const auto OffsetIndexToLargeFor32Bit = Is32Bit && (A.Offset <= -16384 || A.Offset >= 16384);
@@ -103,7 +103,7 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexType = MemOffsetType::SXTX,
.IndexScale = 1,
};
}
@@ -111,15 +111,17 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
if (AtomicTSO) {
// TODO: LRCPC3 support for vector Imm9.
} else if (!Is32Bit && A.Base && (A.Index || A.Segment) && !A.Offset && (A.IndexScale == 1 || A.IndexScale == AccessSizeAsImm)) {
AddressMode B = A;
// ScaledRegisterLoadstore
if (A.Index && A.Segment) {
A.Base = IREmit->Add(GPRSize, A.Base, A.Segment);
} else if (A.Segment) {
A.Index = A.Segment;
A.IndexScale = 1;
if (B.Index && B.Segment) {
B.Base = IREmit->Add(GPRSize, B.Base, B.Segment);
} else if (B.Segment) {
B.Index = B.Segment;
B.IndexScale = 1;
}
return A;
return B;
}
if (Vector || !AtomicTSO) {
@@ -134,7 +136,7 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexType = MemOffsetType::SXTX,
.IndexScale = 1,
};
}
+7 -6
View File
@@ -11,17 +11,18 @@ struct AddressMode {
Ref Segment {nullptr};
Ref Base {nullptr};
Ref Index {nullptr};
MemOffsetType IndexType = MEM_OFFSET_SXTX;
uint8_t IndexScale = 1;
int64_t Offset = 0;
MemOffsetType IndexType = MemOffsetType::SXTX;
uint8_t IndexScale = 1;
// Size in bytes for the address calculation. 8 for an arm64 hardware mode.
IR::OpSize AddrSize;
bool NonTSO;
};
Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage = false);
AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO, bool Vector,
IR::OpSize AccessSize);
Ref LoadEffectiveAddress(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage = false);
AddressMode SelectAddressMode(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO,
bool Vector, IR::OpSize AccessSize);
}; // namespace FEXCore::IR
} // namespace FEXCore::IR
@@ -1,10 +1,10 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "FEXCore/Core/X86Enums.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
@@ -1,30 +1,31 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/JIT/Relocations.h"
#ifdef VIXL_DISASSEMBLER
#include <aarch64/disasm-aarch64.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/vector.h>
#endif
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#include <aarch64/simulator-constants-aarch64.h>
#endif
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/vector.h>
#include <CodeEmitter/Emitter.h>
#include <CodeEmitter/Registers.h>
#include <cstddef>
#include <cstdint>
#include <optional>
#include <span>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::X86State {
enum X86Reg : uint32_t;
}
namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
@@ -104,9 +105,12 @@ constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public ARMEmitter::Emitter {
protected:
public:
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
protected:
FEXCore::Context::ContextImpl* EmitterCTX;
std::span<const ARMEmitter::Register> StaticRegisters {};
@@ -116,8 +120,6 @@ protected:
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
// Correlate an ARM register back to an x86 register index.
+10 -5
View File
@@ -1,14 +1,17 @@
// SPDX-License-Identifier: MIT
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/PrctlUtils.h>
#include <cstdint>
#include "LookupCache.h"
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#endif
@@ -357,6 +360,8 @@ namespace CPU {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
FEXCore::Allocator::VirtualName("FEXMemJIT", reinterpret_cast<void*>(Ptr), Size);
LookupCache = fextl::make_unique<GuestToHostMap>();
}
@@ -395,7 +400,7 @@ namespace CPU {
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(*Buffer);
OnCodeBufferAllocated(Buffer);
return Buffer;
}
+2 -2
View File
@@ -81,7 +81,7 @@ namespace CPU {
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(CodeBuffer&) {};
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
@@ -161,7 +161,7 @@ namespace CPU {
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) = 0;
virtual void ClearCache() {}
+43 -40
View File
@@ -14,6 +14,7 @@ $end_info$
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Syscalls.h>
@@ -88,6 +89,7 @@ namespace ProductNames {
static const char ARM_Blizzard_M2Pro[] = "Apple Blizzard (M2 Pro)";
static const char ARM_Avalanche_M2Max[] = "Apple Avalanche (M2 Max)";
static const char ARM_Blizzard_M2Max[] = "Apple Blizzard (M2 Max)";
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_Ampere_1[] = "AmpereOne";
@@ -188,6 +190,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x61, 0x029, 1, ProductNames::ARM_Firestorm_M1Max}, // Apple Firestorm (M1 Max)
{0x61, 0x025, 1, ProductNames::ARM_Firestorm_M1Pro}, // Apple Firestorm (M1 Pro)
{0x61, 0x023, 1, ProductNames::ARM_Firestorm_M1}, // Apple Firestorm (M1)
{0x61, 0, 1, ProductNames::ARM_AppleSilicon}, // QEmu Apple Silicon
{0x41, 0xd8c, 1, ProductNames::ARM_C1Ultra}, // C1-Ultra
{0x41, 0xd90, 1, ProductNames::ARM_C1Premium}, // C1-Premium
@@ -441,10 +444,10 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
Res.eax = FAMILY_IDENTIFIER;
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(0 << 24); // Local APIC ID
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(GetCPUID() << 24); // Local APIC ID
Res.ecx = (1 << 0) | // SSE3
(CTX->HostFeatures.SupportsPMULL_128Bit << 1) | // PCLMULQDQ
@@ -507,7 +510,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
(1 << 25) | // SSE
(1 << 26) | // SSE2
(0 << 27) | // Self Snoop
(1 << 28) | // Max APIC IDs reserved field is valid
(0 << 28) | // (HTT) Max APIC IDs reserved field is valid
(1 << 29) | // Thermal monitor
(0 << 30) | // Reserved
(0 << 31); // Pending break enable
@@ -909,38 +912,38 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
Res.eax = FAMILY_IDENTIFIER;
Res.ecx = (1 << 0) | // LAHF/SAHF
(1 << 1) | // 0 = Single core product, 1 = multi core product
(0 << 2) | // SVM
(1 << 3) | // Extended APIC register space
(0 << 4) | // LOCK MOV CR0 means MOV CR8
(1 << 5) | // ABM instructions
(0 << 6) | // SSE4a
(0 << 7) | // Misaligned SSE mode
(1 << 8) | // PREFETCHW
(0 << 9) | // OS visible workaround support
(0 << 10) | // Instruction based sampling support
(0 << 11) | // XOP
(0 << 12) | // SKINIT
(0 << 13) | // Watchdog timer support
(0 << 14) | // Reserved
(0 << 15) | // Lightweight profiling support
(0 << 16) | // FMA4
(1 << 17) | // Translation cache extension
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(0 << 20) | // Reserved
(0 << 21) | // XOP-TBM
(0 << 22) | // Topology extensions support
(0 << 23) | // Core performance counter extensions
(0 << 24) | // NB performance counter extensions
(0 << 25) | // Reserved
(0 << 26) | // Data breakpoints extensions
(0 << 27) | // Performance TSC
(0 << 28) | // L2 perf counter extensions
(0 << 29) | // MONITORX
(0 << 30) | // Reserved
(0 << 31); // Reserved
Res.ecx = (1 << 0) | // LAHF/SAHF
(1 << 1) | // 0 = Single core product, 1 = multi core product
(0 << 2) | // SVM
(1 << 3) | // Extended APIC register space
(0 << 4) | // LOCK MOV CR0 means MOV CR8
(1 << 5) | // ABM instructions
(CTX->HostFeatures.SupportsSSE4a << 6) | // SSE4a
(0 << 7) | // Misaligned SSE mode
(1 << 8) | // PREFETCHW
(0 << 9) | // OS visible workaround support
(0 << 10) | // Instruction based sampling support
(0 << 11) | // XOP
(0 << 12) | // SKINIT
(0 << 13) | // Watchdog timer support
(0 << 14) | // Reserved
(0 << 15) | // Lightweight profiling support
(0 << 16) | // FMA4
(1 << 17) | // Translation cache extension
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(0 << 20) | // Reserved
(0 << 21) | // XOP-TBM
(0 << 22) | // Topology extensions support
(0 << 23) | // Core performance counter extensions
(0 << 24) | // NB performance counter extensions
(0 << 25) | // Reserved
(0 << 26) | // Data breakpoints extensions
(0 << 27) | // Performance TSC
(0 << 28) | // L2 perf counter extensions
(0 << 29) | // MONITORX
(0 << 30) | // Reserved
(0 << 31); // Reserved
Res.edx = (1 << 0) | // FPU
(1 << 1) | // Virtual mode extensions
@@ -1094,9 +1097,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) con
(CTX->HostFeatures.SupportsCLZERO << 0); // CLZERO support
uint32_t CoreCount = Cores - 1;
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
((uint32_t)std::log2(CoreCount + 1) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
(std::bit_ceil(Cores) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
return Res;
}
+1 -1
View File
@@ -277,7 +277,7 @@ private:
// 0: Highest function parameter and ID
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 1: Processor info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 2: Cache and TLB info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 3: Serial Number(previously), now reserved
+347 -2
View File
@@ -1,12 +1,213 @@
// SPDX-License-Identifier: MIT
#include <Interface/Context/Context.h>
#include "Utils/SpinWaitLock.h"
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
#include <Interface/Core/JIT/Relocations.h>
#include <Interface/Core/LookupCache.h>
#include <FEXCore/Core/Thunks.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <git_version.h>
#include <xxhash.h>
#include <fstream>
namespace FEXCore {
#if __clang_major__ < 16
ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map, uint64_t FileId, fextl::string Filename)
: SourcecodeMap(std::move(Map))
, FileId(FileId)
, Filename(Filename) {}
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
std::string_view base_filename = FHU::Filesystem::GetFilename(std::string_view {MainExecutable.Filename});
if (FileId != 0xffff'ffff'ffff'ffff) {
return fextl::fmt::format("{}-{:016x}{}", base_filename, MainExecutable.FileId, AddNombSuffix ? "-nomb" : "");
}
return "";
}
fextl::map<CodeMapFileId, CodeMap::ParsedContents> CodeMap::ParseCodeMap(std::ifstream& File) {
fextl::map<CodeMapFileId, CodeMap::ParsedContents> Ret;
while (true) {
Entry Entry;
File.read(reinterpret_cast<char*>(&Entry), sizeof(Entry));
if (!File) {
break;
}
if (Entry.FileId == LoadExternalLibrary.FileId && Entry.BlockOffset == LoadExternalLibrary.BlockOffset) {
ExternalLibraryInfo Info;
File.read(reinterpret_cast<char*>(&Info), sizeof(Info));
fextl::string Filename;
std::getline(File, Filename, '\0');
// Align to 4-byte boundary
char Null[4];
File.read(Null, AlignUp(Filename.size() + 1, 4) - Filename.size() - 1);
if (!File) {
break;
}
Ret[Info.ExternalFileId].Filename = std::move(Filename);
} else if (Entry.FileId == SetExecutableFileId {}.Marker.FileId && Entry.BlockOffset == SetExecutableFileId {}.Marker.BlockOffset) {
CodeMapFileId ExecutableFileId;
File.read(reinterpret_cast<char*>(&ExecutableFileId), sizeof(ExecutableFileId));
if (!File) {
break;
}
Ret[ExecutableFileId].IsExecutable = true;
} else {
if (!Ret.contains(Entry.FileId)) {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
} else {
Ret[Entry.FileId].Blocks.insert(Entry.BlockOffset);
}
}
if (!File) {
break;
}
}
return Ret;
}
CodeMapWriter::CodeMapWriter(CodeMapOpener& Opener, bool OpenEagerly)
: Buffer(4096)
, FileOpener(Opener) {
if (OpenEagerly) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
}
CodeMapWriter::~CodeMapWriter() {
if (CodeMapFD.value_or(-1) != -1) {
Flush(BufferOffset);
close(*CodeMapFD);
}
}
bool CodeMapWriter::IsWriteEnabled(const ExecutableFileSectionInfo& Section) {
if (CodeMapFD == -1) {
return false;
}
// PV libraries can't yet be read by FEXServer, so skip dumping them
if (Section.FileInfo.Filename.starts_with("/run/pressure-vessel")) {
return false;
}
if (CodeMapFD) {
return true;
}
// Acquire mutex and re-check CodeMapFD to avoid race conditions
auto lk = std::unique_lock {Mutex};
if (!CodeMapFD) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
return CodeMapFD != -1;
}
void CodeMapWriter::Flush(size_t Offset) {
// Acquire exclusive lock and flush circular buffer
std::unique_lock Lock {Mutex};
Flush(Offset, Lock);
}
void CodeMapWriter::Flush(size_t Offset, std::unique_lock<std::shared_mutex>&) {
write(*CodeMapFD, Buffer.data(), Offset);
BufferOffset = 0;
}
void CodeMapWriter::AppendBlock(const FEXCore::ExecutableFileSectionInfo& SectionInfo, uint64_t BlockEntry) {
if (!IsWriteEnabled(SectionInfo)) {
return;
}
BlockEntry -= SectionInfo.FileStartVA;
if (BlockEntry > std::numeric_limits<uint32_t>::max()) {
ERROR_AND_DIE_FMT("Cannot write code map");
}
// Register new library if not already known
bool NewLibraryLoad = false;
{
// Check prior registration with shared lock
std::shared_lock Lock {Mutex};
NewLibraryLoad = !KnownFileIds.contains(SectionInfo.FileInfo.FileId);
}
if (NewLibraryLoad) {
// Register to map with exclusive lock
std::unique_lock Lock {Mutex};
NewLibraryLoad &= KnownFileIds.insert(SectionInfo.FileInfo.FileId).second;
}
if (NewLibraryLoad) {
// Add entry to code map
AppendLibraryLoad(SectionInfo.FileInfo);
}
// Register the actual code block
CodeMap::Entry DataEntry {SectionInfo.FileInfo.FileId, static_cast<uint32_t>(BlockEntry)};
AppendData(std::as_bytes(std::span {&DataEntry, 1}));
}
void CodeMapWriter::AppendLibraryLoad(const FEXCore::ExecutableFileInfo& FileInfo) {
// See CodeMap::ExternalLibraryInfo
auto ExternalFileId = FileInfo.FileId;
auto TotalSize = AlignUp(sizeof(CodeMap::LoadExternalLibrary) + sizeof(ExternalFileId) + FileInfo.Filename.size() + 1, 4);
const auto Data = reinterpret_cast<char*>(alloca(TotalSize));
auto WritePtr = std::copy_n(reinterpret_cast<const char*>(&CodeMap::LoadExternalLibrary), sizeof(CodeMap::LoadExternalLibrary), Data);
WritePtr = std::copy_n(reinterpret_cast<const char*>(&ExternalFileId), sizeof(ExternalFileId), WritePtr);
WritePtr = std::copy(FileInfo.Filename.begin(), FileInfo.Filename.end(), WritePtr);
std::fill(WritePtr, Data + TotalSize, 0);
AppendData(std::as_bytes(std::span {Data, TotalSize}));
}
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo) {
CodeMap::SetExecutableFileId Data {.ExecutableFileId = FileInfo.FileId};
AppendData(std::span {reinterpret_cast<const std::byte*>(&Data), sizeof(Data)});
}
void CodeMapWriter::AppendData(std::span<const std::byte> Data) {
std::shared_lock Lock {Mutex};
auto Offset = BufferOffset.fetch_add(Data.size_bytes());
if (Offset + Data.size_bytes() > Buffer.size()) {
// Acquire exclusive lock and flush the buffer.
// Under heavy pressure, multiple threads may observe an exhausted buffer simultaneously.
// The thread with the last in-bounds Offset is responsible for flushing the buffer.
Lock.unlock();
bool IsResponsibleForFlush = false;
{
std::unique_lock ExclusiveLock {Mutex};
IsResponsibleForFlush = (Offset <= Buffer.size());
if (IsResponsibleForFlush) {
Flush(Offset, ExclusiveLock);
}
}
if (!IsResponsibleForFlush) {
// Wait for the buffer to be flushed on the responsible thread
Utils::SpinWaitLock::WaitPred<std::less_equal<>, size_t>(reinterpret_cast<size_t*>(&BufferOffset), Buffer.size());
}
AppendData(Data);
return;
}
memcpy(&Buffer.at(Offset), Data.data(), Data.size_bytes());
}
} // namespace FEXCore
namespace FEXCore::Context {
@@ -15,12 +216,156 @@ CodeCache::CodeCache(ContextImpl& CTX_)
: CTX(CTX_) {}
CodeCache::~CodeCache() = default;
uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
if (Filename.empty()) {
return 0xffff'ffff'ffff'ffff;
}
// For now, we just use the file path as an identifier.
// TODO: Ensure the hash is unique enough to distinguish executables while remaining independent of the installation location
return XXH3_64bits(Filename.data(), Filename.size());
}
struct CodeCacheHeader {
char Magic[4] = {'F', 'X', 'C', 'C'};
uint32_t FormatVersion = 1;
char FEXVersion[8] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
uint32_t CodeBufferSize;
uint32_t NumRelocations;
uint64_t SerializedBaseAddress;
// TODO: Consider including information from LookupCache.BlockLinks
};
void CodeCache::LoadData(Core::InternalThreadState& Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& GuestRIPLookup) {
// TODO
}
template<typename T>
static constexpr auto IsOrderedContainer(const T&) -> std::false_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::map<T...>&) -> std::true_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::set<T...>&) -> std::true_type;
bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const ExecutableFileSectionInfo& SourceBinary, uint64_t SerializedBaseAddress) {
// TODO
auto CodeBuffer = CTX.GetLatest();
auto& LookupCache = *Thread.LookupCache->Shared;
auto Relocations = Thread.CPUBackend->TakeRelocations(SourceBinary.FileStartVA);
// Write file header
CodeCacheHeader header;
memcpy(&header.FEXVersion[0], GIT_SHORT_HASH, strlen(GIT_SHORT_HASH));
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
// Dump guest<->host block mappings
{
// Cache contents must be deterministic, so copy the unordered block list and then sort by key
static_assert(!decltype(IsOrderedContainer(LookupCache.BlockList))::value, "Already deterministic; drop temporary container");
fextl::vector<std::pair<uint64_t, const GuestToHostMap::BlockEntry*>> BlockList;
BlockList.reserve(LookupCache.BlockList.size());
for (auto& [Guest, BlockEntry] : LookupCache.BlockList) {
static_assert(sizeof(Guest) == 8, "Breaking change in code cache data layout");
BlockList.emplace_back(Guest, &BlockEntry);
}
std::ranges::sort(BlockList);
for (auto [Guest, Host] : BlockList) {
static_assert(sizeof(Host->HostCode) == 8, "Breaking change in code cache data layout");
static_assert(sizeof(Host->CodePages[0]) == 8, "Breaking change in code cache data layout");
Guest -= SourceBinary.FileStartVA;
::write(fd, &Guest, sizeof(Guest));
uint64_t HostCode = Host->HostCode - reinterpret_cast<uintptr_t>(CodeBuffer->Ptr);
::write(fd, &HostCode, sizeof(HostCode));
uint64_t NumCodePages = Host->CodePages.size();
::write(fd, &NumCodePages, sizeof(NumCodePages));
LOGMAN_THROW_A_FMT(std::ranges::is_sorted(Host->CodePages), "Code pages aren't sorted");
for (auto CodePage : Host->CodePages) {
CodePage -= SourceBinary.FileStartVA;
::write(fd, &CodePage, sizeof(CodePage));
}
}
}
// Dump relocations
static_assert(sizeof(Relocations[0]) == 48, "Breaking change in code cache data layout");
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
}
// Dump the host code (relocated for position-independent serialization)
std::vector CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Dump code pages
static_assert(decltype(IsOrderedContainer(LookupCache.CodePages))::value, "Non-deterministic data source");
for (auto& [Page, Entrypoints] : LookupCache.CodePages) {
static_assert(sizeof(Page) == 8, "Breaking change in code cache data layout");
::write(fd, &Page, sizeof(Page));
uint64_t NumEntrypoints = Entrypoints.size();
::write(fd, &NumEntrypoints, sizeof(NumEntrypoints));
::write(fd, Entrypoints.data(), Entrypoints.size() * sizeof(Entrypoints[0]));
}
return true;
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
// Generate a literal so we can place it
uint64_t Pointer = ForStorage ? 0 : GetNamedSymbolLiteral(CTX, Reloc.NamedSymbolLiteral.Symbol);
Emitter.dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = ForStorage ? 0 : reinterpret_cast<uint64_t>(CTX.ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Emitter.dc64(GuestEntry + Reloc.GuestRIP.GuestRIP);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
uint64_t Pointer = Reloc.GuestRIP.GuestRIP + GuestEntry;
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIP.RegisterIndex), Pointer, true);
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type {}", ToUnderlying(Reloc.Header.Type));
}
}
return true;
}
+66 -49
View File
@@ -376,7 +376,9 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(Thread);
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>();
Thread->CurrentFrame->Pointers.Common.L1Pointer = Thread->LookupCache->GetL1Pointer();
Thread->CurrentFrame->State.L1Pointer = Thread->LookupCache->GetL1Pointer();
Thread->CurrentFrame->State.L1Mask = Thread->LookupCache->GetScaledL1PointerMask();
Thread->CurrentFrame->Pointers.Common.L2Pointer = Thread->LookupCache->GetPagePointer();
Dispatcher->InitThreadPointers(Thread);
@@ -398,6 +400,7 @@ ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXC
FEXCore::Core::InternalThreadState* Thread = new FEXCore::Core::InternalThreadState {
.CTX = this,
};
FEXCore::Allocator::VirtualName("FEXMem_ThreadState", Thread, sizeof(*Thread));
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = StackPointer;
Thread->CurrentFrame->State.rip = InitialRIP;
@@ -434,6 +437,10 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
Profiler::PostForkAction(Child);
if (Child) {
if (CodeMapWriter) {
CodeMapWriter->ResetAfterFork();
}
CodeInvalidationMutex.StealAndDropActiveLocks();
if (Config.StrictInProcessSplitLocks) {
StrictSplitLockMutex = 0;
@@ -456,9 +463,14 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
}
#endif
void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
void ContextImpl::OnCodeBufferAllocated(const fextl::shared_ptr<CPU::CodeBuffer>& Buffer) {
if (Config.GlobalJITNaming()) {
Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
Symbols.RegisterJITSpace(Buffer->Ptr, Buffer->Size);
}
{
std::scoped_lock lk {CodeBufferListLock};
CodeBufferList.emplace_back(Buffer);
}
}
@@ -470,7 +482,8 @@ void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, boo
Thread->CPUBackend->ClearCache();
} else {
// Clear L1+L2 cache of this thread, and clear L3 cache across any threads using it
Thread->LookupCache->ClearCache();
auto lk = Thread->LookupCache->AcquireWriteLock();
Thread->LookupCache->ClearCache(lk);
}
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
@@ -642,10 +655,10 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
LogMan::Msg::EFmt("Invalid or Unknown instruction: {} 0x{:x}", TableInfo->Name ?: "UND", Block.Entry - GuestRIP);
}
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::NOEXEC_INST) {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
} else {
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::INVALID_INST) {
Thread->OpDispatcher->InvalidOp(DecodedInfo);
} else {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
}
}
@@ -711,7 +724,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (SourcecodeResolver && Config.GDBSymbols()) {
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (MappedSection) {
MappedSection->FileInfo.SourcecodeMap = SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, MappedSection->FileInfo.FileId);
MappedSection->FileInfo.SourcecodeMap =
SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, CodeMap::GetBaseFilename(MappedSection->FileInfo, false));
}
}
@@ -728,7 +742,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
// but this would increase lock contention. Redundant frontend runs aren't
// as expensive and are easily reverted.
if (MaxInst != 1) {
if (auto Block = Thread->LookupCache->FindBlock(GuestRIP)) {
if (auto Block = Thread->LookupCache->FindBlock(Thread, GuestRIP)) {
Thread->OpDispatcher->DelayedDisownBuffer();
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
@@ -769,10 +783,13 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
// Is the code in the cache?
// The backends only check L1 and L2, not L3
if (auto HostCode = Thread->LookupCache->FindBlock(GuestRIP)) {
if (auto HostCode = Thread->LookupCache->FindBlock(Thread, GuestRIP)) {
return HostCode;
}
// Accumulate a JIT count now, as even if another thread raced us, it should count as a compile.
FEXCORE_PROFILE_INSTANT_INCREMENT(Thread, AccumulatedJITCount, 1);
auto [CompiledCode, DebugData, StartAddr, Length, NeedsAddGuestCodeRanges] = CompileCode(Thread, GuestRIP, MaxInst);
auto CodePtr = CompiledCode.EntryPoints[GuestRIP];
if (CodePtr == nullptr) {
@@ -826,20 +843,32 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
Thread->CPUBackend->ClearRelocations();
}
fextl::vector<uint64_t> CodePages;
if (NeedsAddGuestCodeRanges) {
// Track in the guest to host map all entrypoints for all pages the compiled block touches, if any page didn't previously
// contain code, inform the frontend so it can setup SMC detection.
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
CodePages.reserve(BlockInfo->CodePages.size());
CodePages.insert(CodePages.end(), BlockInfo->CodePages.begin(), BlockInfo->CodePages.end());
for (auto CodePage : BlockInfo->CodePages) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
if (Thread->LookupCache->AddBlockExecutableRange(Thread, BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
// Insert to lookup cache
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(GuestAddr, HostAddr);
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, CodePages, HostAddr);
}
if (CodeMapWriter) {
auto Region = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (Region && Region->FileStartVA != 0) {
CodeMapWriter->AppendBlock(*Region, GuestRIP);
}
}
return (uintptr_t)CodePtr;
@@ -866,49 +895,37 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
void ContextImpl::InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) {
FEXCORE_PROFILE_SCOPED("InvalidateCodeBuffersCodeRange");
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::scoped_lock lk {CodeBufferListLock};
auto it = CodeBufferList.begin();
while (it != CodeBufferList.end()) {
if (auto Strong = it->lock()) {
Strong->LookupCache->InvalidateRange(Start, Length);
it++;
} else {
it = CodeBufferList.erase(it);
}
}
}
void ContextImpl::InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
auto lk = Thread->LookupCache->AcquireLock();
auto& CodePages = Thread->LookupCache->Shared->CodePages;
if (Thread->LookupCache->InvalidateCacheRange(Start, Length)) {
FEXCORE_PROFILE_SCOPED("InvalidateCallRet");
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
Accumulator.emplace_back(std::move(it->second));
}
bool InvalidatedAnyEntries = false;
for (const auto& PageEntries : Accumulator) {
for (const auto& Entry : PageEntries) {
if (ContextImpl::ThreadRemoveCodeEntry(Thread, Entry)) {
InvalidatedAnyEntries = true;
}
}
}
if (InvalidatedAnyEntries) {
// This may cause access violations in the thread on Windows as zeroing is not atomic, this is handled by the frontend
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
}
void ContextImpl::ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
static_cast<ContextImpl*>(Frame->Thread->CTX)->SyscallHandler->InvalidateGuestCodeRange(Frame->Thread, GuestRIP, 1);
}
@@ -952,10 +969,10 @@ void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t Gu
if (GPRSize == IR::OpSize::i64Bit) {
IR::Ref R = emit->_StoreRegister(emit->Constant(Entrypoint), GPRSize);
R->Reg = IR::PhysicalRegister(IR::GPRFixedClass, X86State::REG_R11).Raw;
R->Reg = IR::PhysicalRegister(IR::RegClass::GPRFixed, X86State::REG_R11).Raw;
} else {
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
emit->_StoreContextFPR(GPRSize, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
}
emit->_ExitFunction(IR::OpSize::i64Bit, emit->Constant(GuestThunkEntrypoint), IR::BranchHint::None, emit->Invalid(), emit->Invalid());
},
@@ -976,7 +993,7 @@ void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t Gu
void ContextImpl::AddForceTSOInformation(const IntervalList<uint64_t>& ValidRanges, fextl::set<uint64_t>&& Instructions) {
LogMan::Throw::AFmt(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
ForceTSOValidRanges.Insert(ValidRanges);
ForceTSOInstructions.merge(Instructions);
ForceTSOInstructions.merge(std::move(Instructions));
}
void ContextImpl::RemoveForceTSOInformation(uint64_t Address, uint64_t Size) {
@@ -5,7 +5,6 @@
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Config/Config.h>
@@ -36,12 +35,14 @@ static void SleepThread(FEXCore::Context::ContextImpl* CTX, FEXCore::Core::CpuSt
CTX->SyscallHandler->SleepThread(CTX, Frame);
}
constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 4;
constexpr size_t MAX_DISPATCHER_CODE_SIZE = FEXCore::Utils::FEX_PAGE_SIZE * 4;
Dispatcher::Dispatcher(FEXCore::Context::ContextImpl* ctx)
: Arm64Emitter(ctx, FEXCore::Allocator::VirtualAlloc(MAX_DISPATCHER_CODE_SIZE, true), MAX_DISPATCHER_CODE_SIZE)
, CTX {ctx} {
EmitDispatcher();
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(GetBufferBase()), MAX_DISPATCHER_CODE_SIZE);
}
Dispatcher::~Dispatcher() {
@@ -96,7 +97,7 @@ void Dispatcher::EmitDispatcher() {
ARMEmitter::BiDirectionalLabel LoopTop {};
#ifdef _M_ARM_64EC
b(&LoopTop);
(void)b(&LoopTop);
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
@@ -104,10 +105,10 @@ void Dispatcher::EmitDispatcher() {
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
// Force a single instruction block if ENTRY_FILL_SRA_SINGLE_INST_REG is nonzero entering the JIT, used for inline SMC handling.
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
(void)cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
// Enter JIT
b(&LoopTop);
(void)b(&LoopTop);
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
// Load ThreadState and write the target PC there
@@ -128,7 +129,7 @@ void Dispatcher::EmitDispatcher() {
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, REG_CALLRET_SP);
// EC_CALL_CHECKER_PC_REG is REG_PF which isn't touched by any of the above
sub(ARMEmitter::Size::i64Bit, TMP1, EC_CALL_CHECKER_PC_REG, TMP1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
// If the entry at the TOS is for the target address, pop it and return to the JIT code
add(ARMEmitter::Size::i64Bit, REG_CALLRET_SP, REG_CALLRET_SP, 0x10);
@@ -167,67 +168,73 @@ void Dispatcher::EmitDispatcher() {
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
(void)!Bind(&l_NotECCode);
(void)Bind(&l_NotECCode);
#endif
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
(void)cbnz(ARMEmitter::Size::i32Bit, TMP1, &CompileSingleStep);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
if (std::popcount(VirtualMemorySize) == 1) {
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), VirtualMemorySize - 1);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, VirtualMemorySize);
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), TMP4);
}
ARMEmitter::ForwardLabel NoBlock;
{
// Offset the address and add to our page pointer
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP4, 12);
if (DisableL2Cache()) {
(void)b(&NoBlock);
} else {
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Load the pointer from the offset
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 3);
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
if (std::popcount(VirtualMemorySize) == 1) {
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), VirtualMemorySize - 1);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, VirtualMemorySize);
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), TMP4);
}
// If page pointer is zero then we have no block
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
// Steal the page offset
and_(ARMEmitter::Size::i64Bit, TMP2, TMP4, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// The the full LookupCacheEntry with a single LDP.
// Check the guest address first to ensure it maps to the address we are currently at.
// This fixes aliasing problems
ldp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP2, TMP1, 0);
// If the guest address doesn't match, Compile the block.
sub(TMP2, TMP2, RipReg);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP2, &NoBlock);
// Check the host address to see if it matches, else compile the block.
(void)cbz(ARMEmitter::Size::i64Bit, TMP4, &NoBlock);
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
// Offset the address and add to our page pointer
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP4, 12);
and_(ARMEmitter::Size::i64Bit, TMP2, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, 4);
stp<ARMEmitter::IndexType::OFFSET>(TMP4, RipReg, TMP1);
// Load the pointer from the offset
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 3);
// Jump to the block
br(TMP4);
// If page pointer is zero then we have no block
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
// Steal the page offset
and_(ARMEmitter::Size::i64Bit, TMP2, TMP4, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(sizeof(LookupCache::LookupCacheEntry)));
// The the full LookupCacheEntry with a single LDP.
// Check the guest address first to ensure it maps to the address we are currently at.
// This fixes aliasing problems
ldp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP2, TMP1, 0);
// If the guest address doesn't match, Compile the block.
sub(TMP2, TMP2, RipReg);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP2, &NoBlock);
// Check the host address to see if it matches, else compile the block.
(void)cbz(ARMEmitter::Size::i64Bit, TMP4, &NoBlock);
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.L1Pointer));
// Calculate (tmp1 + ((ripreg & L1_ENTRIES_MASK) << 4)) for the address
// L1Mask is pre-shifted.
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, RipReg.R(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(sizeof(LookupCache::LookupCacheEntry)));
add(TMP1, TMP1, TMP2);
stp<ARMEmitter::IndexType::OFFSET>(TMP4, RipReg, TMP1);
// Jump to the block
br(TMP4);
}
}
}
@@ -481,7 +488,7 @@ void Dispatcher::EmitDispatcher() {
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->X86CodeGen.CallbackReturn);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->SignalDelegation->GetThunkCallbackRET());
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, CTX->Config.Is64BitMode ? 16 : 12);
@@ -4,6 +4,7 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <array>
@@ -50,6 +51,10 @@ public:
}
#endif
uint64_t GetExitFunctionLinkerAddress() const {
return ExitFunctionLinkerAddress;
}
SignalDelegatorConfig MakeSignalDelegatorConfig() const;
protected:
@@ -92,6 +97,8 @@ private:
void EmitDispatcher();
uint64_t GenerateABICall(FallbackABI ABI);
FEX_CONFIG_OPT(DisableL2Cache, DISABLEL2CACHE);
};
} // namespace FEXCore::CPU
+8 -8
View File
@@ -9,7 +9,6 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/LookupCache.h"
#include <array>
@@ -90,11 +89,6 @@ Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
}
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
if (EntryPoint == CTX->X86CodeGen.CallbackReturn) {
return true;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
@@ -1047,8 +1041,11 @@ Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
auto Result = ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
DecodedBlockStatus::NOEXEC_INST;
DecodeInst->InstSize = 0;
return ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST : DecodedBlockStatus::NOEXEC_INST;
return Result;
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
return DecodedBlockStatus::INVALID_INST;
@@ -1450,7 +1447,10 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" : "NoExec", OpAddress);
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
"PartialDecode",
OpAddress);
}
break;
}
+1
View File
@@ -27,6 +27,7 @@ public:
SUCCESS,
INVALID_INST,
NOEXEC_INST,
PARTIAL_DECODE_INST,
};
// New Frontend decoding
+25 -3
View File
@@ -50,14 +50,13 @@ DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
auto Dst = GetReg(Node);
uint64_t Mask = ~0ULL;
const auto OpSize = IROp->Size;
if (OpSize == IR::OpSize::i32Bit) {
Mask = 0xFFFF'FFFFULL;
}
LoadConstant(ARMEmitter::Size::i64Bit, Dst, Constant & Mask);
InsertGuestRIPMove(GetReg(Node), Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -372,7 +371,7 @@ DEF_OP(CondSubNZCV) {
DEF_OP(Neg) {
auto Op = IROp->C<IR::IROp_Neg>();
if (Op->Cond == FEXCore::IR::COND_AL) {
if (Op->Cond == IR::CondClass::AL) {
neg(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src));
} else {
cneg(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src), MapCC(Op->Cond));
@@ -1190,6 +1189,19 @@ DEF_OP(Rev) {
}
}
DEF_OP(Rbit) {
auto Op = IROp->C<IR::IROp_Rbit>();
const auto OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = ConvertSize48(IROp);
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src);
rbit(EmitSize, Dst, Src);
}
DEF_OP(Bfi) {
auto Op = IROp->C<IR::IROp_Bfi>();
const auto EmitSize = ConvertSize(IROp);
@@ -1273,6 +1285,16 @@ DEF_OP(Sbfe) {
sbfx(ConvertSize(IROp), Dst, Src, Op->lsb, Op->Width);
}
DEF_OP(MaskGenerateFromBitWidth) {
auto Op = IROp->C<IR::IROp_MaskGenerateFromBitWidth>();
auto BitWidth = GetReg(Op->BitWidth);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, -1);
cmp(ARMEmitter::Size::i64Bit, BitWidth, 0);
lslv(ARMEmitter::Size::i64Bit, TMP2, TMP1, BitWidth);
csinv(ARMEmitter::Size::i64Bit, GetReg(Node), TMP1, TMP2, ARMEmitter::Condition::CC_EQ);
}
DEF_OP(Select) {
auto Op = IROp->C<IR::IROp_Select>();
const auto OpSize = IROp->Size;
@@ -11,23 +11,18 @@ $end_info$
#include <FEXCore/Core/Thunks.h>
namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl& CTX, FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op)); break;
return CTX.Dispatcher->GetExitFunctionLinkerAddress();
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum& Sum) {
Relocation MoveABI {};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE};
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
@@ -38,9 +33,9 @@ void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR
}
Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
uint64_t Pointer = GetNamedSymbolLiteral(*CTX, Op);
Arm64JITCore::NamedSymbolLiteralPair Lit {
NamedSymbolLiteralPair Lit {
.Lit = Pointer,
.MoveABI =
{
@@ -48,92 +43,72 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
{
.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit) {
switch (Lit.MoveABI.Header.Type) {
case RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL:
case RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Lit.MoveABI.Header.Offset = GetCursorOffset();
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type for {}", __FUNCTION__);
}
BindOrRestart(&Lit.Loc);
dc64(Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
auto Arm64JITCore::InsertGuestRIPLiteral(uint64_t GuestRIP) -> NamedSymbolLiteralPair {
return {
.Lit = GuestRIP,
.MoveABI =
{
.GuestRIP = {.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL,
},
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
.GuestRIP = GuestRIP},
},
};
}
void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant) {
Relocation MoveABI {};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
MoveABI.GuestRIP.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE};
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
MoveABI.GuestRIP.GuestRIP = Constant;
MoveABI.GuestRIP.RegisterIndex = Reg.Idx();
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, false);
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation> Relocations) {
const auto OrigBase = GetBufferBase();
const auto OrigSize = GetBufferSize();
const auto OrigOffset = GetCursorOffset();
SetBuffer(reinterpret_cast<std::uint8_t*>(Code.data()), Code.size_bytes());
for (auto& Reloc : Relocations) {
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc.NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(Reloc.NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIPMove.RegisterIndex), Pointer, true);
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations(uint64_t GuestBaseAddress) {
// Rebase relocations to library base address
for (auto& Relocation : Relocations) {
switch (Relocation.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE:
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Relocation.GuestRIP.GuestRIP -= GuestBaseAddress;
break;
}
default:;
}
}
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return true;
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations() {
return std::move(Relocations);
}
@@ -138,26 +138,6 @@ DEF_OP(CAS) {
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
eor(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
(void)cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
const auto OpSize = IROp->Size;
+21 -125
View File
@@ -174,14 +174,12 @@ DEF_OP(ExitFunction) {
}
// L1 Cache
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer));
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.L1Pointer));
// Calculate (tmp1 + ((ripreg & L1_ENTRIES_MASK) << 4)) for the address
// arithmetic. ubfiz+add is marginally faster on Firestorm than
// and+add(shift). Same performance on Cortex.
static_assert(LookupCache::L1_ENTRIES_MASK == ((1u << 20) - 1));
ubfiz(ARMEmitter::Size::i64Bit, TMP4, RipReg, 4, 20);
add(TMP1, TMP1, TMP4);
// L1Mask is pre-shifted.
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, RipReg, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(sizeof(LookupCache::LookupCacheEntry)));
add(TMP1, TMP1, TMP2);
ldp<ARMEmitter::IndexType::OFFSET>(TMP2, TMP1, TMP1, 0);
@@ -235,16 +233,16 @@ DEF_OP(CondJump) {
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1), "CondJump: Expected GPR");
LOGMAN_THROW_A_FMT(isConst, "CondJump: Expected constant source");
if (Op->Cond.Val == FEXCore::IR::COND_EQ) {
if (Op->Cond == IR::CondClass::EQ) {
LOGMAN_THROW_A_FMT(Const == 0, "CondJump: Expected 0 source");
cbz_OrRestart(Size, Reg, TrueTargetLabel);
} else if (Op->Cond.Val == FEXCore::IR::COND_NEQ) {
} else if (Op->Cond == IR::CondClass::NEQ) {
LOGMAN_THROW_A_FMT(Const == 0, "CondJump: Expected 0 source");
cbnz_OrRestart(Size, Reg, TrueTargetLabel);
} else if (Op->Cond.Val == FEXCore::IR::COND_TSTZ) {
} else if (Op->Cond == IR::CondClass::TSTZ) {
LOGMAN_THROW_A_FMT(Const < 64, "CondJump: Expected valid bit source");
tbz_OrRestart(Reg, Const, TrueTargetLabel);
} else if (Op->Cond.Val == FEXCore::IR::COND_TSTNZ) {
} else if (Op->Cond == IR::CondClass::TSTNZ) {
LOGMAN_THROW_A_FMT(Const < 64, "CondJump: Expected valid bit source");
tbnz_OrRestart(Reg, Const, TrueTargetLabel);
} else {
@@ -262,16 +260,10 @@ DEF_OP(Syscall) {
// X1: ThreadState
// X2: Pointer to SyscallArguments
FEXCore::IR::SyscallFlags Flags = Op->Flags;
PushDynamicRegs(TMP1);
uint32_t GPRSpillMask = ~0U;
uint32_t FPRSpillMask = ~0U;
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) == FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) {
// Need to spill all caller saved registers still
GPRSpillMask = CALLER_GPR_MASK;
FPRSpillMask = CALLER_FPR_MASK;
}
SpillStaticRegs(TMP1, true, GPRSpillMask, FPRSpillMask);
@@ -305,117 +297,22 @@ DEF_OP(Syscall) {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
PopDynamicRegs();
PopDynamicRegs();
if ((Flags & FEXCore::IR::SyscallFlags::NORETURNEDRESULT) != FEXCore::IR::SyscallFlags::NORETURNEDRESULT) {
// Move result to its destination register.
// Only if `NORETURNEDRESULT` wasn't set, otherwise we might overwrite the CPUState refilled with `FillStaticRegs`
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
}
}
}
const auto OSABI = CTX->SyscallHandler->GetOSABI();
DEF_OP(InlineSyscall) {
auto Op = IROp->C<IR::IROp_InlineSyscall>();
// Arguments are passed as follows:
// X8: SyscallNumber - RA INTERSECT
// X0: Arg0 & Return
// X1: Arg1
// X2: Arg2
// X3: Arg3
// X4: Arg4 - RA INTERSECT
// X5: Arg5 - RA INTERSECT
// X6: Arg6 - Doesn't exist in x86-64 land. RA INTERSECT
// One argument is removed from the SyscallArguments::MAX_ARGS since the first argument was syscall number
const static std::array<ARMEmitter::XRegister, FEXCore::HLE::SyscallArguments::MAX_ARGS - 1> RegArgs = {
{ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::XReg::x2, ARMEmitter::XReg::x3, ARMEmitter::XReg::x4, ARMEmitter::XReg::x5}};
bool Intersects {};
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS - 1; ++i) {
if (Op->Header.Args[i].IsInvalid()) {
break;
}
auto Reg = GetReg(Op->Header.Args[i]);
if (Reg == ARMEmitter::Reg::r8 || Reg == ARMEmitter::Reg::r4 || Reg == ARMEmitter::Reg::r5) {
SpillMask |= (1U << Reg.Idx());
Intersects = true;
}
}
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(TMP1, false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
// 16bit LoadConstant to be a single instruction
// We must always spill at least one register (x8) so this value always has a bit set
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SpillMask & 0xFFFF);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
// Now that we have claimed to be a syscall we can set up the arguments
const auto EmitSize = CTX->Config.Is64BitMode() ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto EmitSubSize = CTX->Config.Is64BitMode() ? ARMEmitter::SubRegSize::i64Bit : ARMEmitter::SubRegSize::i32Bit;
if (Intersects) {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS - 1; ++i) {
if (Op->Header.Args[i].IsInvalid()) {
break;
}
auto Reg = GetReg(Op->Header.Args[i]);
if (SpillMask & (1U << Reg.Idx())) {
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RDX, and RSP. Which have just been spilled
// Just load back from the context.
auto Correlation = GetX86RegRelationToARMReg(Reg);
LOGMAN_THROW_A_FMT(Correlation != X86State::REG_INVALID, "Invalid register mapping");
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[Correlation]));
} else {
mov(EmitSize, RegArgs[i].R(), Reg);
}
}
} else {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS - 1; ++i) {
if (Op->Header.Args[i].IsInvalid()) {
break;
}
mov(EmitSize, RegArgs[i].R(), GetReg(Op->Header.Args[i]));
}
}
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, Op->HostSyscallNumber);
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
if ((Op->Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r1);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
// Result is now in x0
// Move result to its destination register
mov(EmitSize, GetReg(Node), ARMEmitter::Reg::r0);
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_GENERIC) {
// Move result to its destination register.
// Only if `NORETURNEDRESULT` wasn't set, otherwise we might overwrite the CPUState refilled with `FillStaticRegs`
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
}
}
@@ -431,8 +328,7 @@ DEF_OP(Thunk) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr));
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
InsertNamedThunkRelocation(ARMEmitter::Reg::r2, Op->ThunkNameHash);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
} else {
@@ -423,11 +423,11 @@ DEF_OP(Vector_FToI) {
const auto Mask = PRED_TMP_32B.Merging();
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Host.Val: frinti(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::Nearest: frintn(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::NegInfinity: frintm(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::PosInfinity: frintp(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::TowardsZero: frintz(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::Host: frinti(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
}
} else {
const auto IsScalar = ElementSize == OpSize;
@@ -449,21 +449,21 @@ DEF_OP(Vector_FToI) {
}
switch (Op->Round) {
case IR::Round_Nearest.Val: ROUNDING_FN(frintn); break;
case IR::Round_Negative_Infinity.Val: ROUNDING_FN(frintm); break;
case IR::Round_Positive_Infinity.Val: ROUNDING_FN(frintp); break;
case IR::Round_Towards_Zero.Val: ROUNDING_FN(frintz); break;
case IR::Round_Host.Val: ROUNDING_FN(frinti); break;
case IR::RoundMode::Nearest: ROUNDING_FN(frintn); break;
case IR::RoundMode::NegInfinity: ROUNDING_FN(frintm); break;
case IR::RoundMode::PosInfinity: ROUNDING_FN(frintp); break;
case IR::RoundMode::TowardsZero: ROUNDING_FN(frintz); break;
case IR::RoundMode::Host: ROUNDING_FN(frinti); break;
}
#undef ROUNDING_FN
} else {
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Host.Val: frinti(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Nearest: frintn(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::NegInfinity: frintm(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::PosInfinity: frintp(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::TowardsZero: frintz(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Host: frinti(SubEmitSize, Dst.Q(), Vector.Q()); break;
}
}
}
@@ -539,11 +539,11 @@ DEF_OP(Vector_F64ToI32) {
// Then convert to integers using fcvtzs.
auto CVTReg = Dst.Z();
switch (Round) {
case IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Towards_Zero.Val: CVTReg = Vector.Z(); break;
case IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::Nearest: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::NegInfinity: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::PosInfinity: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::TowardsZero: CVTReg = Vector.Z(); break;
case IR::RoundMode::Host: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
}
fcvtzs(Dst.Z(), ARMEmitter::SubRegSize::i32Bit, Mask, CVTReg, ARMEmitter::SubRegSize::i64Bit);
@@ -567,11 +567,11 @@ DEF_OP(Vector_F64ToI32) {
///< Round float to integral depending on rounding mode.
switch (Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Nearest: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::NegInfinity: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::PosInfinity: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::TowardsZero: frintz(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Host: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
}
// Now narrow from f64 to f32.
+76 -53
View File
@@ -493,7 +493,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
static void DirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = JumpThunkStartAddress / 4 - CallerAddress / 4;
@@ -511,11 +511,12 @@ static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Co
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
static void IndirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
BranchEmit.b(0x8);
// Restore branch +2 instructions to jump to the linker block
BranchEmit.b(0x2);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
@@ -538,7 +539,7 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
// Guard the LookupCache lock with the code invalidation mutex, to avoid issues with forking
auto lk_inval =
GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
HostCode = Thread->LookupCache->FindBlock(GuestRip);
HostCode = Thread->LookupCache->FindBlock(Thread, GuestRip);
}
if (!HostCode) {
// Hold a reference to the code buffer, to avoid linking unmapped code if compilation triggers a recreation.
@@ -563,7 +564,7 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
auto lk_inval = GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
// Lock here is necessary to prevent simultaneous linking and delinking
auto lk = Thread->LookupCache->AcquireLock();
auto lk = Thread->LookupCache->AcquireWriteLock();
// For non-calls, this would extend into the block's code, however that's fine as an out-of-range adr would never
// be generated avoiding any false positives.
@@ -576,14 +577,12 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
if (KnownCallMarkerInst == ExpectedKnownCallMarkerInst) {
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(GuestRip, Record, [](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, true);
});
Thread->LookupCache->AddBlockLink(
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, true); }, lk);
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(GuestRip, Record, [](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, false);
});
Thread->LookupCache->AddBlockLink(
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, false); }, lk);
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
@@ -602,7 +601,7 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(LdrInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
Thread->LookupCache->AddBlockLink(GuestRip, Record, IndirectBlockDelinker);
Thread->LookupCache->AddBlockLink(GuestRip, Record, IndirectBlockDelinker, lk);
}
return HostCode;
@@ -623,10 +622,10 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
RAPass->AddRegisters(FEXCore::IR::GPRClass, GeneralRegisters.size());
RAPass->AddRegisters(FEXCore::IR::GPRFixedClass, StaticRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRClass, GeneralFPRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRFixedClass, StaticFPRegisters.size());
RAPass->AddRegisters(IR::RegClass::GPR, GeneralRegisters.size());
RAPass->AddRegisters(IR::RegClass::GPRFixed, StaticRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPR, GeneralFPRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPRFixed, StaticFPRegisters.size());
RAPass->PairRegs = PairRegisters;
{
@@ -678,13 +677,13 @@ void Arm64JITCore::EmitDetectionString() {
void Arm64JITCore::ClearCache() {
// NOTE: Holding on to the reference here is required to ensure validity of the WriteLock mutex
auto PrevCodeBuffer = CurrentCodeBuffer;
std::lock_guard lk(PrevCodeBuffer->LookupCache->WriteLock);
auto lk = PrevCodeBuffer->LookupCache->AcquireWriteLock();
auto CodeBuffer = GetEmptyCodeBuffer();
SetBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
EmitDetectionString();
ThreadState->LookupCache->ChangeGuestToHostMapping(*PrevCodeBuffer, *CurrentCodeBuffer->LookupCache);
ThreadState->LookupCache->ChangeGuestToHostMapping(*PrevCodeBuffer, *CurrentCodeBuffer->LookupCache, lk);
}
Arm64JITCore::~Arm64JITCore() {}
@@ -784,7 +783,7 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
ldr(TMP2.W(), STATE_PTR(CpuStateFrame, SuspendDoorbell));
ARMEmitter::ForwardLabel l_NoSuspend;
cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
(void)cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
brk(SuspendMagic);
(void)Bind(&l_NoSuspend);
#endif
@@ -810,22 +809,35 @@ void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool C
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
EmitSuspendInterruptCheck();
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
const auto PrevNumAllocations = Relocations.size();
this->Entry = Entry;
this->DebugData = DebugData;
this->IR = IR;
RequiresFarARM64Jumps = false;
SSANodeMultiplier = 24;
switch (static_cast<RestartOptions::Control>(FEXCore::LongJump::SetJump(RestartControl.RestartJump))) {
// Prepare restart via long jump in case branch encoding fails.
// This uses UncheckedLongJump since we don't implement std::longjmp in WoA setups
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(ThreadState->RestartJump))) {
case RestartOptions::Control::Incoming:
// Nothing
break;
case RestartOptions::Control::EnableFarARM64Jumps: RequiresFarARM64Jumps = true; break;
default: ERROR_AND_DIE_FMT("Unhandled Arm64 restart condition!");
case RestartOptions::Control::NeedsLargerJITSpace:
// Get rid of the claimed buffer immediately, we can't fit in it at all.
TempAllocator.UnclaimBuffer();
SSANodeMultiplier *= 2;
break;
default: LOGMAN_MSG_A_FMT("Unhandled Arm64 restart condition!");
}
uint32_t SSACount = IR->GetSSACount();
@@ -837,12 +849,19 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CodeData.EntryPoints.clear();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = 0x1000 + SSACount * 24;
// One page baseline, plus SSANodeMultipler bytes, plus another page for guard page.
const uint32_t DesiredBufferRange = AlignUp(FEXCore::Utils::FEX_PAGE_SIZE * 2 + SSACount * SSANodeMultiplier, FEXCore::Utils::FEX_PAGE_SIZE);
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBuffer = TempAllocator.ReownOrClaimBuffer(BufferRange);
SetBuffer(TempCodeBuffer, BufferRange);
auto TempCodeBufferInfo = TempAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBuffer = TempCodeBufferInfo.Ptr;
const uint32_t UsableBufferRange = TempCodeBufferInfo.Size - FEXCore::Utils::FEX_PAGE_SIZE;
SetBuffer(TempCodeBuffer, UsableBufferRange);
ThreadState->JITGuardPage = reinterpret_cast<uintptr_t>(TempCodeBuffer) + UsableBufferRange;
ThreadState->JITGuardOverflowArgument = FEXCore::ToUnderlying(RestartOptions::Control::NeedsLargerJITSpace);
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
@@ -932,8 +951,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
switch (IROp->Op) {
#define REGISTER_OP_RT(op, x) \
case FEXCore::IR::IROps::OP_##op: std::invoke(RT_##x, this, IROp, CodeNode); break
#define REGISTER_OP(op, x) \
case FEXCore::IR::IROps::OP_##op: Op_##x(IROp, CodeNode); break
@@ -974,22 +991,28 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// This is a ExitFunctionLinkData struct
BindOrRestart(&l_ExitLink);
dc64(0); // HostCode
dc64(PendingJumpThunk.GuestRIP); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
dc64(0); // HostCode
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(PendingJumpThunk.GuestRIP)); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
BindOrRestart(&l_ExitLink);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
PlaceNamedSymbolLiteral(InsertNamedSymbolLiteral(RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER));
// CodeSize not including the header or tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeBegin;
// Add the JitCodeTail
// Add the JitCodeTail (written later)
Align(alignof(JITCodeTail));
auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
const auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
JITCodeTail JITBlockTail {
.RIP = Entry,
.GuestSize = Size,
.SpinLockFutex = 0,
.SingleInst = SingleInst,
};
// Entries that live after the JITCodeTail.
// These entries correlate JIT code regions with guest RIP regions.
@@ -1007,23 +1030,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// FEXCore::Utils::vl64 GuestRIPOffset;
// };
auto JITRIPEntriesBegin = GetCursorAddress<uint8_t*>();
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
JITBlockTail->GuestSize = Size;
JITBlockTail->SingleInst = SingleInst;
JITBlockTail->SpinLockFutex = 0;
const auto JITRIPEntriesBegin = JITBlockTailLocation + sizeof(JITBlockTail);
auto JITRIPEntriesLocation = JITRIPEntriesBegin;
{
// Store the RIP entries.
JITBlockTail->NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail->OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
JITBlockTail.NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail.OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
uintptr_t CurrentRIPOffset = 0;
uint64_t CurrentPCOffset = 0;
@@ -1039,14 +1052,20 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
}
CursorIncrement(JITRIPEntriesLocation - JITRIPEntriesBegin);
SetCursorOffset(JITRIPEntriesLocation - CodeData.BlockBegin);
Align();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
// Finalize and write block tail data
JITBlockTail.Size = CodeData.Size;
{
auto PrevCur = GetCursorOffset();
memcpy(JITBlockTailLocation, &JITBlockTail, sizeof(JITBlockTail));
SetCursorOffset(JITBlockTailLocation - CodeData.BlockBegin + offsetof(JITCodeTail, RIP));
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(JITBlockTail.RIP));
SetCursorOffset(PrevCur);
}
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
@@ -1055,7 +1074,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Query size of generated code
const auto TempSize = GetCursorOffset();
LOGMAN_THROW_A_FMT(TempSize <= BufferRange, "Exceeded bounds of temporary buffer ({:#x} vs {:#x})", TempSize, BufferRange);
// Bring CodeBuffer up to date
{
@@ -1063,7 +1081,8 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
"doesn't match up!\n");
if (auto Prev = CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(ThreadState->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
ThreadState->LookupCache->ChangeGuestToHostMapping(*Prev, *CurrentCodeBuffer->LookupCache);
auto lk = ThreadState->LookupCache->AcquireWriteLock();
ThreadState->LookupCache->ChangeGuestToHostMapping(*Prev, *CurrentCodeBuffer->LookupCache, lk);
}
// NOTE: 16-byte alignment of the new cursor offset must be preserved for block linking records
@@ -1086,6 +1105,10 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
CodeBegin += Delta;
for (std::size_t Idx = PrevNumAllocations; Idx != Relocations.size(); ++Idx) {
Relocations[Idx].Header.Offset += CodeBuffers.LatestOffset;
}
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
+122 -110
View File
@@ -10,13 +10,17 @@ $end_info$
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/Relocations.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IntrusiveIRList.h"
#include "Interface/IR/RegisterAllocationData.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/LongJump.h>
@@ -26,16 +30,19 @@ $end_info$
#include <array>
#include <cstdint>
#include <functional>
#include <optional>
#include <utility>
#include <variant>
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::Context {
struct ExitFunctionLinkData;
}
namespace FEXCore::IR {
class RegisterAllocationPass;
}
namespace FEXCore::CPU {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
@@ -54,9 +61,6 @@ public:
}
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(HalfBarrierTSOEnabled, HALFBARRIERTSOENABLED);
const bool HostSupportsSVE128 {};
const bool HostSupportsSVE256 {};
const bool HostSupportsAVX256 {};
@@ -64,10 +68,10 @@ private:
const bool HostSupportsAFP {};
struct RestartOptions {
FEXCore::LongJump::JumpBuf RestartJump;
enum class Control : uint64_t {
Incoming = 0,
EnableFarARM64Jumps = 1,
NeedsLargerJITSpace = 2,
};
};
@@ -75,6 +79,8 @@ private:
// In the rare case when those assumptions are broken, FEX needs to safely restart the JIT.
RestartOptions RestartControl {};
bool RequiresFarARM64Jumps {};
// Default to 6 instructions per SSA node.
uint32_t SSANodeMultiplier {24};
ARMEmitter::BiDirectionalLabel* PendingTargetLabel {};
ARMEmitter::BiDirectionalLabel* PendingCallReturnTargetLabel {};
@@ -105,11 +111,13 @@ private:
[[nodiscard]]
ARMEmitter::Register GetReg(IR::PhysicalRegister Reg) const {
LOGMAN_THROW_A_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
const auto RegClass = Reg.AsRegClass();
if (Reg.Class == IR::GPRFixedClass.Val) {
LOGMAN_THROW_A_FMT(RegClass == IR::RegClass::GPRFixed || RegClass == IR::RegClass::GPR, "Unexpected Class: {}", Reg.Class);
if (RegClass == IR::RegClass::GPRFixed) {
return StaticRegisters[Reg.Reg];
} else if (Reg.Class == IR::GPRClass.Val) {
} else if (RegClass == IR::RegClass::GPR) {
return GeneralRegisters[Reg.Reg];
}
@@ -128,11 +136,13 @@ private:
[[nodiscard]]
ARMEmitter::VRegister GetVReg(IR::PhysicalRegister Reg) const {
LOGMAN_THROW_A_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
const auto RegClass = Reg.AsRegClass();
if (Reg.Class == IR::FPRFixedClass.Val) {
LOGMAN_THROW_A_FMT(RegClass == IR::RegClass::FPRFixed || RegClass == IR::RegClass::FPR, "Unexpected Class: {}", Reg.Class);
if (RegClass == IR::RegClass::FPRFixed) {
return StaticFPRegisters[Reg.Reg];
} else if (Reg.Class == IR::FPRClass.Val) {
} else if (RegClass == IR::RegClass::FPR) {
return GeneralFPRegisters[Reg.Reg];
}
@@ -150,8 +160,8 @@ private:
}
[[nodiscard]]
FEXCore::IR::RegisterClassType GetRegClass(IR::Ref Node) const {
return FEXCore::IR::RegisterClassType {IR::PhysicalRegister(Node).Class};
static IR::RegClass GetRegClass(IR::Ref Node) {
return IR::PhysicalRegister(Node).AsRegClass();
}
[[nodiscard]]
@@ -168,7 +178,7 @@ private:
// Converts IR-base shift type to ARMEmitter shift type.
// Will be a no-op, only a type conversion since the two definitions match.
[[nodiscard]]
ARMEmitter::ShiftType ConvertIRShiftType(IR::ShiftType Shift) const {
static ARMEmitter::ShiftType ConvertIRShiftType(IR::ShiftType Shift) {
return Shift == IR::ShiftType::LSL ? ARMEmitter::ShiftType::LSL :
Shift == IR::ShiftType::LSR ? ARMEmitter::ShiftType::LSR :
Shift == IR::ShiftType::ASR ? ARMEmitter::ShiftType::ASR :
@@ -176,18 +186,23 @@ private:
}
[[nodiscard]]
ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
static ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
return Op->Size == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
}
[[nodiscard]]
ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
static ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->Size == IR::OpSize::i32Bit || Op->Size == IR::OpSize::i64Bit, "Invalid size");
return ConvertSize(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(IR::OpSize ElementSize) {
static ARMEmitter::Size ConvertSize(IR::OpSize Size) {
return Size == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
}
[[nodiscard]]
static ARMEmitter::SubRegSize ConvertSubRegSize16(IR::OpSize ElementSize) {
LOGMAN_THROW_A_FMT(ElementSize == IR::OpSize::i8Bit || ElementSize == IR::OpSize::i16Bit || ElementSize == IR::OpSize::i32Bit ||
ElementSize == IR::OpSize::i64Bit || ElementSize == IR::OpSize::i128Bit,
"Invalid size");
@@ -199,105 +214,105 @@ private:
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize16(const IR::IROp_Header* Op) {
return ConvertSubRegSize16(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(IR::OpSize ElementSize) {
static ARMEmitter::SubRegSize ConvertSubRegSize8(IR::OpSize ElementSize) {
LOGMAN_THROW_A_FMT(ElementSize != IR::OpSize::i128Bit, "Invalid size");
return ConvertSubRegSize16(ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize8(const IR::IROp_Header* Op) {
return ConvertSubRegSize8(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i64Bit, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i8Bit, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair16(const IR::IROp_Header* Op) {
static ARMEmitter::VectorRegSizePair ConvertSubRegSizePair16(const IR::IROp_Header* Op) {
return ARMEmitter::ToVectorSizePair(ConvertSubRegSize16(Op));
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
static ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i128Bit, "Invalid size");
return ConvertSubRegSizePair16(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
static ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i8Bit, "Invalid size");
return ConvertSubRegSizePair8(Op);
}
[[nodiscard]]
ARMEmitter::Condition MapCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU:
case FEXCore::IR::COND_VS: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU:
case FEXCore::IR::COND_VC: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
static ARMEmitter::Condition MapCC(IR::CondClass Cond) {
switch (Cond) {
case IR::CondClass::EQ: return ARMEmitter::Condition::CC_EQ;
case IR::CondClass::NEQ: return ARMEmitter::Condition::CC_NE;
case IR::CondClass::SGE: return ARMEmitter::Condition::CC_GE;
case IR::CondClass::SLT: return ARMEmitter::Condition::CC_LT;
case IR::CondClass::SGT: return ARMEmitter::Condition::CC_GT;
case IR::CondClass::SLE: return ARMEmitter::Condition::CC_LE;
case IR::CondClass::UGE: return ARMEmitter::Condition::CC_CS;
case IR::CondClass::ULT: return ARMEmitter::Condition::CC_CC;
case IR::CondClass::UGT: return ARMEmitter::Condition::CC_HI;
case IR::CondClass::ULE: return ARMEmitter::Condition::CC_LS;
case IR::CondClass::FLU: return ARMEmitter::Condition::CC_LT;
case IR::CondClass::FGE: return ARMEmitter::Condition::CC_GE;
case IR::CondClass::FLEU: return ARMEmitter::Condition::CC_LE;
case IR::CondClass::FGT: return ARMEmitter::Condition::CC_GT;
case IR::CondClass::FU:
case IR::CondClass::VS: return ARMEmitter::Condition::CC_VS;
case IR::CondClass::FNU:
case IR::CondClass::VC: return ARMEmitter::Condition::CC_VC;
case IR::CondClass::MI: return ARMEmitter::Condition::CC_MI;
case IR::CondClass::PL: return ARMEmitter::Condition::CC_PL;
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
}
}
[[nodiscard]]
bool IsFPR(IR::RegisterClassType Class) const {
return Class == IR::FPRClass || Class == IR::FPRFixedClass;
static bool IsFPR(IR::RegClass Class) {
return Class == IR::RegClass::FPR || Class == IR::RegClass::FPRFixed;
}
[[nodiscard]]
bool IsGPR(IR::RegisterClassType Class) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
static bool IsGPR(IR::RegClass Class) {
return Class == IR::RegClass::GPR || Class == IR::RegClass::GPRFixed;
}
[[nodiscard]]
bool IsGPR(IR::Ref Node) {
static bool IsGPR(IR::Ref Node) {
return IsGPR(GetRegClass(Node));
}
[[nodiscard]]
bool IsFPR(IR::Ref Node) {
static bool IsFPR(IR::Ref Node) {
return IsFPR(GetRegClass(Node));
}
[[nodiscard]]
bool IsGPR(IR::OrderedNodeWrapper Wrap) {
return IsGPR(IR::RegisterClassType {IR::PhysicalRegister(Wrap).Class});
static bool IsGPR(IR::OrderedNodeWrapper Wrap) {
return IsGPR(IR::PhysicalRegister(Wrap).AsRegClass());
}
[[nodiscard]]
bool IsFPR(IR::OrderedNodeWrapper Wrap) {
return IsFPR(IR::RegisterClassType {IR::PhysicalRegister(Wrap).Class});
static bool IsFPR(IR::OrderedNodeWrapper Wrap) {
return IsFPR(IR::PhysicalRegister(Wrap).AsRegClass());
}
[[nodiscard]]
@@ -339,9 +354,7 @@ private:
}
// Restart helpers
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void bl_OrRestart(T* Label) {
if (bl(Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
@@ -349,12 +362,10 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void b_OrRestart(T* Label) {
if (b(Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
@@ -362,12 +373,10 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void b_OrRestart(ARMEmitter::Condition Cond, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
@@ -385,12 +394,10 @@ private:
return;
}
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void cbz_OrRestart(ARMEmitter::Size s, ARMEmitter::Register rt, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
@@ -408,12 +415,10 @@ private:
return;
}
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void cbnz_OrRestart(ARMEmitter::Size s, ARMEmitter::Register rt, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
@@ -431,12 +436,10 @@ private:
return;
}
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void tbz_OrRestart(ARMEmitter::Register rt, uint32_t Bit, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
@@ -454,12 +457,10 @@ private:
return;
}
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void tbnz_OrRestart(ARMEmitter::Register rt, uint32_t Bit, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
@@ -477,38 +478,40 @@ private:
return;
}
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void adr_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADR.");
}
return;
}
if (adr(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADR currently unsupported!");
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void adrp_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADRP.");
}
return;
}
if (adrp(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADRP currently unsupported!");
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<typename T>
requires (std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>)
template<ARMEmitter::IsLabel T>
void BindOrRestart(T* Label) {
if (Bind(Label)) {
return;
@@ -516,11 +519,11 @@ private:
if (RequiresFarARM64Jumps) {
// This should have been caught before this point.
ERROR_AND_DIE_FMT("Oops. Unhandled long bind.");
ERROR_AND_DIE_FMT("Unhandled long bind");
return;
}
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
@@ -533,8 +536,6 @@ private:
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
@@ -571,19 +572,30 @@ private:
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Inserts a relocation for a constant value relative to the guest entrypoint
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
NamedSymbolLiteralPair InsertGuestRIPLiteral(uint64_t GuestRIP);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit);
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit);
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation>);
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() override;
/**
* Returns any relocations generated since the last call to TakeRelocations.
*
* GuestBaseAddress must match the base virtual address to which the
* input x86 binary is mapped.
*/
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) override;
/** @} */
@@ -606,7 +618,7 @@ private:
void Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize, ARMEmitter::VRegister Dst, ARMEmitter::VRegister IncomingDst,
std::optional<ARMEmitter::Register> BaseAddr, ARMEmitter::VRegister VectorIndexLow,
std::optional<ARMEmitter::VRegister> VectorIndexHigh, ARMEmitter::VRegister MaskReg, IR::OpSize VectorIndexSize,
size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale);
size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale, IR::OpSize AddrSize);
void EmitTFCheck();
+72 -84
View File
@@ -21,7 +21,7 @@ DEF_OP(LoadContext) {
const auto Op = IROp->C<IR::IROp_LoadContext>();
const auto OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
auto Dst = GetReg(Node);
switch (OpSize) {
@@ -52,7 +52,7 @@ DEF_OP(LoadContext) {
DEF_OP(LoadContextPair) {
const auto Op = IROp->C<IR::IROp_LoadContextPair>();
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst1 = GetReg(Op->OutValue1);
const auto Dst2 = GetReg(Op->OutValue2);
@@ -78,7 +78,7 @@ DEF_OP(StoreContext) {
const auto Op = IROp->C<IR::IROp_StoreContext>();
const auto OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
auto Src = GetZeroableReg(Op->Value);
switch (OpSize) {
@@ -110,7 +110,7 @@ DEF_OP(StoreContextPair) {
const auto Op = IROp->C<IR::IROp_StoreContextPair>();
const auto OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
auto Src1 = GetZeroableReg(Op->Value1);
auto Src2 = GetZeroableReg(Op->Value2);
@@ -135,11 +135,11 @@ DEF_OP(StoreContextPair) {
DEF_OP(LoadRegister) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
if (Op->Class == IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
LOGMAN_THROW_A_FMT(Op->Reg < StaticRegisters.size(), "out of range reg");
mov(GetReg(Node).X(), StaticRegisters[Op->Reg].X());
} else if (Op->Class == IR::FPRClass) {
} else if (Op->Class == IR::RegClass::FPR) {
const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Op->Reg < StaticFPRegisters.size(), "out of range reg");
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
@@ -175,12 +175,13 @@ DEF_OP(LoadAF) {
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
auto Reg = IR::PhysicalRegister(Node);
const auto Reg = IR::PhysicalRegister(Node);
const auto RegClass = Reg.AsRegClass();
if (Reg.Class == IR::GPRFixedClass) {
if (RegClass == IR::RegClass::GPRFixed) {
// Always use 64-bit, it's faster. Upper bits ignored for 32-bit mode.
mov(ARMEmitter::Size::i64Bit, GetReg(Reg), GetReg(Op->Value));
} else if (Reg.Class == IR::FPRFixedClass) {
} else if (RegClass == IR::RegClass::FPRFixed) {
const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
@@ -193,7 +194,7 @@ DEF_OP(StoreRegister) {
mov(guest.Q(), host.Q());
}
} else {
LOGMAN_THROW_A_FMT(false, "Unhandled Op->Class {}", Reg.Class);
LOGMAN_THROW_A_FMT(false, "Unhandled Op->Class {}", RegClass);
}
}
@@ -225,7 +226,7 @@ DEF_OP(LoadContextIndexed) {
const auto Index = GetReg(Op->Index);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
switch (Op->Stride) {
case 1:
case 2:
@@ -288,7 +289,7 @@ DEF_OP(StoreContextIndexed) {
const auto Index = GetReg(Op->Index);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Value = GetReg(Op->Value);
switch (Op->Stride) {
@@ -372,7 +373,7 @@ DEF_OP(SpillRegister) {
const auto OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * MaxSpillSlotSize;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Src = GetReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i8Bit: {
@@ -413,7 +414,7 @@ DEF_OP(SpillRegister) {
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize); break;
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
} else if (Op->Class == FEXCore::IR::RegClass::FPR) {
const auto Src = GetVReg(Op->Value);
switch (OpSize) {
@@ -452,7 +453,7 @@ DEF_OP(SpillRegister) {
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize); break;
}
} else {
LOGMAN_MSG_A_FMT("Unhandled SpillRegister class: {}", Op->Class.Val);
LOGMAN_MSG_A_FMT("Unhandled SpillRegister class: {}", Op->Class);
}
}
@@ -461,7 +462,7 @@ DEF_OP(FillRegister) {
const auto OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * MaxSpillSlotSize;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
switch (OpSize) {
case IR::OpSize::i8Bit: {
@@ -502,7 +503,7 @@ DEF_OP(FillRegister) {
}
default: LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize); break;
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
} else if (Op->Class == FEXCore::IR::RegClass::FPR) {
const auto Dst = GetVReg(Node);
switch (OpSize) {
@@ -541,7 +542,7 @@ DEF_OP(FillRegister) {
default: LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize); break;
}
} else {
LOGMAN_MSG_A_FMT("Unhandled FillRegister class: {}", Op->Class.Val);
LOGMAN_MSG_A_FMT("Unhandled FillRegister class: {}", Op->Class);
}
}
@@ -578,14 +579,14 @@ ARMEmitter::ExtendedMemOperand Arm64JITCore::GenerateMemOperand(
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, Const);
} else {
auto RegOffset = GetReg(Offset);
switch (OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val:
switch (OffsetType) {
case IR::MemOffsetType::SXTX:
return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::SXTX, FEXCore::ilog2(OffsetScale));
case IR::MEM_OFFSET_UXTW.Val:
case IR::MemOffsetType::UXTW:
return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::UXTW, FEXCore::ilog2(OffsetScale));
case IR::MEM_OFFSET_SXTW.Val:
case IR::MemOffsetType::SXTW:
return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
default: LOGMAN_MSG_A_FMT("Unhandled GenerateMemOperand OffsetType: {}", OffsetType.Val); break;
default: LOGMAN_MSG_A_FMT("Unhandled GenerateMemOperand OffsetType: {}", OffsetType); break;
}
}
}
@@ -612,20 +613,20 @@ ARMEmitter::Register Arm64JITCore::ApplyMemOperand(IR::OpSize AccessSize, ARMEmi
add(ARMEmitter::Size::i64Bit, Tmp, Base, Tmp, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
} else {
auto RegOffset = GetReg(Offset);
switch (OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val:
switch (OffsetType) {
case IR::MemOffsetType::SXTX:
add(ARMEmitter::Size::i64Bit, Tmp, Base, RegOffset, ARMEmitter::ExtendedType::SXTX, FEXCore::ilog2(OffsetScale));
break;
case IR::MEM_OFFSET_UXTW.Val:
case IR::MemOffsetType::UXTW:
add(ARMEmitter::Size::i64Bit, Tmp, Base, RegOffset, ARMEmitter::ExtendedType::UXTW, FEXCore::ilog2(OffsetScale));
break;
case IR::MEM_OFFSET_SXTW.Val:
case IR::MemOffsetType::SXTW:
add(ARMEmitter::Size::i64Bit, Tmp, Base, RegOffset, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
break;
default: LOGMAN_MSG_A_FMT("Unhandled OffsetType: {}", OffsetType.Val); break;
default: LOGMAN_MSG_A_FMT("Unhandled OffsetType: {}", OffsetType); break;
}
}
return Tmp;
@@ -676,7 +677,7 @@ ARMEmitter::SVEMemOperand Arm64JITCore::GenerateSVEMemOperand(IR::OpSize AccessS
// Note that we do nothing with the offset type and offset scale,
// since SVE loads and stores don't have the ability to perform an
// optional extension or shift as part of their behavior.
LOGMAN_THROW_A_FMT(OffsetType.Val == IR::MEM_OFFSET_SXTX.Val, "Currently only the default offset type (SXTX) is supported.");
LOGMAN_THROW_A_FMT(OffsetType == IR::MemOffsetType::SXTX, "Currently only the default offset type (SXTX) is supported.");
const auto RegOffset = GetReg(Offset);
return ARMEmitter::SVEMemOperand(Base.X(), RegOffset.X());
@@ -689,7 +690,7 @@ DEF_OP(LoadMem) {
const auto MemReg = GetReg(Op->Addr);
const auto MemSrc = GenerateMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
switch (OpSize) {
@@ -723,7 +724,7 @@ DEF_OP(LoadMemPair) {
const auto Op = IROp->C<IR::IROp_LoadMemPair>();
const auto Addr = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst1 = GetReg(Op->OutValue1);
const auto Dst2 = GetReg(Op->OutValue2);
@@ -751,13 +752,13 @@ DEF_OP(LoadMemTSO) {
const auto MemReg = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid() || CTX->HostFeatures.SupportsTSOImm9, "unexpected offset");
LOGMAN_THROW_A_FMT(Op->OffsetScale == 1, "unexpected offset scale");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MEM_OFFSET_SXTX, "unexpected offset type");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MemOffsetType::SXTX, "unexpected offset type");
}
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
@@ -776,12 +777,10 @@ DEF_OP(LoadMemTSO) {
case IR::OpSize::i64Bit: ldapur(Dst.X(), MemReg, Offset); break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", OpSize); break;
}
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
// Half-barrier once back-patched.
nop();
}
// Half-barrier once back-patched.
nop();
}
} else if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
} else if (CTX->HostFeatures.SupportsRCPC && Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
if (OpSize == IR::OpSize::i8Bit) {
// 8bit load is always aligned to natural alignment
@@ -793,12 +792,10 @@ DEF_OP(LoadMemTSO) {
case IR::OpSize::i64Bit: ldapr(Dst.X(), MemReg); break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", OpSize); break;
}
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
// Half-barrier once back-patched.
nop();
}
// Half-barrier once back-patched.
nop();
}
} else if (Op->Class == FEXCore::IR::GPRClass) {
} else if (Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
if (OpSize == IR::OpSize::i8Bit) {
// 8bit load is always aligned to natural alignment
@@ -810,10 +807,8 @@ DEF_OP(LoadMemTSO) {
case IR::OpSize::i64Bit: ldar(Dst.X(), MemReg); break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", OpSize); break;
}
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
// Half-barrier once back-patched.
nop();
}
// Half-barrier once back-patched.
nop();
}
} else {
const auto Dst = GetVReg(Node);
@@ -1045,7 +1040,7 @@ void Arm64JITCore::Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize,
ARMEmitter::VRegister IncomingDst, std::optional<ARMEmitter::Register> BaseAddr,
ARMEmitter::VRegister VectorIndexLow, std::optional<ARMEmitter::VRegister> VectorIndexHigh,
ARMEmitter::VRegister MaskReg, IR::OpSize VectorIndexSize, size_t DataElementOffsetStart,
size_t IndexElementOffsetStart, uint8_t OffsetScale) {
size_t IndexElementOffsetStart, uint8_t OffsetScale, IR::OpSize AddrSize) {
LOGMAN_THROW_A_FMT(ElementSize >= IR::OpSize::i8Bit && ElementSize <= IR::OpSize::i64Bit, "Invalid element size");
const auto PerformSMove = [this](IR::OpSize ElementSize, const ARMEmitter::Register Dst, const ARMEmitter::VRegister Vector, int index) {
@@ -1121,17 +1116,17 @@ void Arm64JITCore::Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize,
// Calculate memory position for this gather load
if (BaseAddr.has_value()) {
if (VectorIndexSize == IR::OpSize::i32Bit) {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
add(ConvertSize(AddrSize), TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
} else {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
add(ConvertSize(AddrSize), TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
}
} else {
///< In this case we have no base address, All addresses come from the vector register itself
if (VectorIndexSize == IR::OpSize::i32Bit) {
// Sign extend and shift in to the 64-bit register
sbfiz(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale), 32);
sbfiz(ConvertSize(AddrSize), TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale), 32);
} else {
lsl(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale));
lsl(ConvertSize(AddrSize), TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale));
}
}
@@ -1189,7 +1184,8 @@ DEF_OP(VLoadVectorGatherMasked) {
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
const bool SupportsSVELoad = (HostSupportsSVE128 || HostSupportsSVE256) &&
(OffsetScale == 1 || OffsetScale == IR::OpSizeToSize(VectorIndexSize)) && VectorIndexSize == IROp->ElementSize;
(OffsetScale == 1 || OffsetScale == IR::OpSizeToSize(VectorIndexSize)) &&
VectorIndexSize == IROp->ElementSize && Op->AddrSize == IR::OpSize::i64Bit;
if (SupportsSVELoad) {
uint8_t SVEScale = FEXCore::ilog2(OffsetScale);
@@ -1247,7 +1243,7 @@ DEF_OP(VLoadVectorGatherMasked) {
} else {
LOGMAN_THROW_A_FMT(!Is256Bit, "Can't emulate this gather load in the backend! Programming error!");
Emulate128BitGather(IROp->Size, IROp->ElementSize, Dst, IncomingDst, BaseAddr, VectorIndexLow, VectorIndexHigh, MaskReg,
VectorIndexSize, DataElementOffsetStart, IndexElementOffsetStart, OffsetScale);
VectorIndexSize, DataElementOffsetStart, IndexElementOffsetStart, OffsetScale, Op->AddrSize);
}
}
@@ -1272,7 +1268,9 @@ DEF_OP(VLoadVectorGatherMaskedQPS) {
!Op->VectorIndexHigh.IsInvalid() ? std::make_optional(GetVReg(Op->VectorIndexHigh)) : std::nullopt;
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
if (HostSupportsSVE128 && (OffsetScale == 1 || OffsetScale == 4)) {
const bool SupportsSVELoad = HostSupportsSVE128 && (OffsetScale == 1 || OffsetScale == 4) && Op->AddrSize == IR::OpSize::i64Bit;
if (SupportsSVELoad) {
ARMEmitter::SVEModType ModType = ARMEmitter::SVEModType::MOD_NONE;
if (OffsetScale != 1) {
ModType = ARMEmitter::SVEModType::MOD_LSL;
@@ -1326,7 +1324,7 @@ DEF_OP(VLoadVectorGatherMaskedQPS) {
}
} else {
Emulate128BitGather(IR::OpSize::i128Bit, IR::OpSize::i32Bit, Dst, IncomingDst, BaseAddr, VectorIndexLow, VectorIndexHigh, MaskReg,
IR::OpSize::i64Bit, 0, 0, OffsetScale);
IR::OpSize::i64Bit, 0, 0, OffsetScale, Op->AddrSize);
}
}
@@ -1627,7 +1625,7 @@ DEF_OP(StoreMem) {
const auto MemReg = GetReg(Op->Addr);
const auto MemSrc = GenerateMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i8Bit: strb(Src, MemSrc); break;
@@ -1738,7 +1736,7 @@ DEF_OP(StoreMemPair) {
const auto OpSize = IROp->Size;
const auto Addr = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Src1 = GetZeroableReg(Op->Value1);
const auto Src2 = GetZeroableReg(Op->Value2);
switch (OpSize) {
@@ -1765,13 +1763,13 @@ DEF_OP(StoreMemTSO) {
const auto MemReg = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid() || CTX->HostFeatures.SupportsTSOImm9, "unexpected offset");
LOGMAN_THROW_A_FMT(Op->OffsetScale == 1, "unexpected offset scale");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MEM_OFFSET_SXTX, "unexpected offset type");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MemOffsetType::SXTX, "unexpected offset type");
}
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
@@ -1783,10 +1781,8 @@ DEF_OP(StoreMemTSO) {
// 8bit load is always aligned to natural alignment
stlurb(Src, MemReg, Offset);
} else {
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
// Half-barrier once back-patched.
nop();
}
// Half-barrier once back-patched.
nop();
switch (OpSize) {
case IR::OpSize::i16Bit: stlurh(Src, MemReg, Offset); break;
case IR::OpSize::i32Bit: stlur(Src.W(), MemReg, Offset); break;
@@ -1794,17 +1790,15 @@ DEF_OP(StoreMemTSO) {
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", OpSize); break;
}
}
} else if (Op->Class == FEXCore::IR::GPRClass) {
} else if (Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
if (OpSize == IR::OpSize::i8Bit) {
// 8bit load is always aligned to natural alignment
stlrb(Src, MemReg);
} else {
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
// Half-barrier once back-patched.
nop();
}
// Half-barrier once back-patched.
nop();
switch (OpSize) {
case IR::OpSize::i16Bit: stlrh(Src, MemReg); break;
case IR::OpSize::i32Bit: stlr(Src.W(), MemReg); break;
@@ -1898,9 +1892,7 @@ DEF_OP(MemSet) {
// 8bit load is always aligned to natural alignment
stlrb(Value.W(), TMP2);
} else {
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
nop();
}
nop();
switch (OpSize) {
case 2: stlrh(Value.W(), TMP2); break;
case 4: stlr(Value.W(), TMP2); break;
@@ -2118,11 +2110,9 @@ DEF_OP(MemCpy) {
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size); break;
}
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
// Placeholders for backpatching barriers (one per load/store)
nop();
nop();
}
// Placeholders for backpatching barriers (one per load/store)
nop();
nop();
switch (OpSize) {
case 2: stlrh(TMP4.W(), TMP2); break;
@@ -2144,11 +2134,9 @@ DEF_OP(MemCpy) {
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size); break;
}
if (HalfBarrierTSOEnabled() && !ParanoidTSO()) {
// Placeholders for backpatching barriers (one per load/store)
nop();
nop();
}
// Placeholders for backpatching barriers (one per load/store)
nop();
nop();
switch (OpSize) {
case 2: stlrh(TMP4.W(), TMP2); break;
@@ -15,6 +15,7 @@ $end_info$
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/EnumUtils.h>
namespace FEXCore::CPU {
@@ -47,10 +48,10 @@ DEF_OP(GuestOpcode) {
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
case IR::Fence_Load.Val: dmb(ARMEmitter::BarrierScope::LD); break;
case IR::Fence_LoadStore.Val: dmb(ARMEmitter::BarrierScope::SY); break;
case IR::Fence_Store.Val: dmb(ARMEmitter::BarrierScope::ST); break;
case IR::Fence_Inst.Val: isb(); break;
case IR::FenceType::Load: dmb(ARMEmitter::BarrierScope::LD); break;
case IR::FenceType::LoadStore: dmb(ARMEmitter::BarrierScope::SY); break;
case IR::FenceType::Store: dmb(ARMEmitter::BarrierScope::ST); break;
case IR::FenceType::Inst: isb(); break;
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
}
}
@@ -107,10 +108,10 @@ DEF_OP(GetRoundingMode) {
// zero. Just swapping 01 and 10. That's a bitfield reverse. Round mode is in
// bottom two bits. After reversing as a 32-bit operation, it'll be in [31:30]
// and ripe for reinsertion back at 0.
static_assert(IR::ROUND_MODE_NEAREST == 0);
static_assert(IR::ROUND_MODE_NEGATIVE_INFINITY == 1);
static_assert(IR::ROUND_MODE_POSITIVE_INFINITY == 2);
static_assert(IR::ROUND_MODE_TOWARDS_ZERO == 3);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::Nearest) == 0);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::NegInfinity) == 1);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::PosInfinity) == 2);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::TowardsZero) == 3);
rbit(ARMEmitter::Size::i32Bit, TMP1, Dst);
bfi(ARMEmitter::Size::i64Bit, Dst, TMP1, 30, 2);
@@ -18,18 +18,4 @@ DEF_OP(RMWHandle) {
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(IROp->Args[0]));
}
DEF_OP(Swap1) {
auto Op = IROp->C<IR::IROp_Swap1>();
auto A = GetReg(Op->A), B = GetReg(Op->B);
LOGMAN_THROW_A_FMT(B == GetReg(Node), "Invariant");
mov(ARMEmitter::Size::i64Bit, TMP1, A);
mov(ARMEmitter::Size::i64Bit, A, B);
mov(ARMEmitter::Size::i64Bit, B, TMP1);
}
DEF_OP(Swap2) {
// Implemented above
}
} // namespace FEXCore::CPU
+34 -24
View File
@@ -1,79 +1,89 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
enum class RelocationTypes : uint8_t {
enum class RelocationTypes : uint32_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// 4 instruction constant generation
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// 8 byte literal (relative to binary base address)
RELOC_GUEST_RIP_LITERAL,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
// 4 instruction constant generation
// Aligned to struct RelocGuestRIP
RELOC_GUEST_RIP_MOVE,
};
struct RelocationTypeHeader final {
struct FEX_PACKED RelocationHeader final {
// Offset to the relocated host code data
uint64_t Offset {};
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
enum class NamedSymbol : uint32_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
RelocationTypeHeader Header {};
RelocationHeader Header {};
NamedSymbol Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
uint32_t Pad[8];
};
struct RelocNamedThunkMove final {
RelocationTypeHeader Header {};
RelocationHeader Header {};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
uint32_t RegisterIndex;
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
struct RelocGuestRIPMove final {
RelocationTypeHeader Header {};
struct RelocGuestRIP final {
RelocationHeader Header {};
// GPR index the constant is being moved to
// GPR index the constant is being moved to (for non-literal relocations)
uint8_t RegisterIndex;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
char Pad[3];
// The unrelocated RIP that is being moved
// The base RIP (to be moved by the register for non-literal relocations).
// In a serialized code cache, this is relative to the binary base address.
uint64_t GuestRIP;
uint32_t pad2[6] {};
};
union Relocation {
RelocationTypeHeader Header {};
RelocationHeader Header {};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIPMove GuestRIPMove;
RelocGuestRIP GuestRIP;
};
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl&, RelocNamedSymbolLiteral::NamedSymbol);
} // namespace FEXCore::CPU
@@ -41,6 +41,7 @@ namespace FEXCore::CPU {
const auto Op = IROp->C<IR::IROp_##FEXOp>(); \
const auto OpSize = IROp->Size; \
const auto Is256Bit = OpSize == IR::OpSize::i256Bit; \
const auto Is128Bit = OpSize == IR::OpSize::i128Bit; \
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__); \
\
const auto Dst = GetVReg(Node); \
@@ -49,8 +50,10 @@ namespace FEXCore::CPU {
\
if (HostSupportsSVE256 && Is256Bit) { \
ARMOp(Dst.Z(), Vector1.Z(), Vector2.Z()); \
} else { \
} else if (Is128Bit) { \
ARMOp(Dst.Q(), Vector1.Q(), Vector2.Q()); \
} else { \
ARMOp(Dst.D(), Vector1.D(), Vector2.D()); \
} \
}
@@ -744,11 +747,11 @@ DEF_OP(VFToIScalarInsert) {
auto Src = *std::get_if<ARMEmitter::VRegister>(&SrcVar);
switch (RoundMode) {
case IR::Round_Nearest: frintn(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Negative_Infinity: frintm(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Positive_Infinity: frintp(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Towards_Zero: frintz(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Host: frinti(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::Nearest: frintn(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::NegInfinity: frintm(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::PosInfinity: frintp(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::TowardsZero: frintz(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::Host: frinti(SubRegSize.Scalar, Dst, Src); break;
}
};
+24 -17
View File
@@ -15,7 +15,7 @@ $end_info$
namespace FEXCore {
GuestToHostMap::GuestToHostMap()
: BlockLinks_mbr {fextl::pmr::get_default_resource()} {
: BlockLinks_mbr {"FEXMem_BlockLinks"} {
BlockLinks_pma = fextl::make_unique<std::pmr::polymorphic_allocator<std::byte>>(&BlockLinks_mbr);
// Setup our PMR map.
BlockLinks = BlockLinks_pma->new_object<BlockLinksMapType>();
@@ -24,7 +24,7 @@ GuestToHostMap::GuestToHostMap()
LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
: ctx {CTX} {
TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
TotalCacheSize = ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE + MAX_L1_SIZE;
// Block cache ends up looking like this
// PageMemoryMap[VirtualMemoryRegion >> 12]
@@ -39,6 +39,10 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We need one pointer per page of virtual memory
// At 64GB of virtual memory this will allocate 128MB of virtual memory space
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
// Allocate our memory backing our pages
@@ -46,14 +50,21 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// XXX: We can drop down to 16KB if we store 4byte offsets from the code base
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = PagePointer + ctx->Config.VirtualMemSize / 4096 * 8;
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
PageMemory = PagePointer + ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8;
// L1 Cache
L1Pointer = PageMemory + CODE_SIZE;
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
FEXCore::Allocator::VirtualName("FEXMem_Lookup_L1", reinterpret_cast<void*>(L1Pointer), MAX_L1_SIZE);
VirtualMemSize = ctx->Config.VirtualMemSize;
if (DynamicL1Cache()) {
// Start at minimum size when dynamic.
L1PointerMask = MIN_L1_ENTRIES - 1;
} else {
// Start at maximum instead.
L1PointerMask = MAX_L1_ENTRIES - 1;
}
}
LookupCache::~LookupCache() {
@@ -64,31 +75,27 @@ LookupCache::~LookupCache() {
// These will get freed when their memory allocators are deallocated.
}
void LookupCache::ClearL2Cache() {
auto lk = Shared->AcquireLock();
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE, false);
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE, false);
AllocateOffset = 0;
}
void LookupCache::ClearThreadLocalCaches() {
auto lk = Shared->AcquireLock();
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
CachedCodePages.clear();
}
void LookupCache::ClearCache() {
auto lk = Shared->AcquireLock();
void LookupCache::ClearCache(const LookupCacheWriteLockToken& lk) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
ClearThreadLocalCaches(lk);
Shared->ClearCache(lk);
}
void GuestToHostMap::ClearCache(const LockToken&) {
void GuestToHostMap::ClearCache(const LookupCacheWriteLockToken&) {
// Allocate a new pointer from the BlockLinks pma again.
BlockLinks = BlockLinks_pma->new_object<BlockLinksMapType>();
// All code is gone, clear the block list
+267 -121
View File
@@ -2,30 +2,57 @@
#pragma once
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/robin_set.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/memory_resource.h>
#include <cstdint>
#include <functional>
#include <stddef.h>
#include <utility>
#include <mutex>
namespace FEXCore {
struct LookupCacheBaseLockToken {
protected:
// Protected constructor - only derived classes can construct
LookupCacheBaseLockToken() = default;
};
struct LookupCacheWriteLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheWriteLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
: Lock {Mutex} {}
std::lock_guard<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct LookupCacheReadLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheReadLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
: Lock {Mutex} {}
std::shared_lock<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct GuestToHostMap {
std::recursive_mutex WriteLock;
struct LockToken {
std::lock_guard<std::recursive_mutex> Lock;
};
FEXCore::Utils::WritePriorityMutex::Mutex Lock {};
[[nodiscard]]
LockToken AcquireLock() {
return LockToken {std::lock_guard {WriteLock}};
LookupCacheWriteLockToken AcquireWriteLock() {
return LookupCacheWriteLockToken {Lock};
}
[[nodiscard]]
LookupCacheReadLockToken AcquireReadLock() {
return LookupCacheReadLockToken {Lock};
}
struct BlockLinkTag {
@@ -49,53 +76,72 @@ struct GuestToHostMap {
// walking each block member and destructing objects.
//
// This makes `BlockLinks` look like a raw pointer that could memory leak, but since it is backed by the MBR, it won't.
std::pmr::monotonic_buffer_resource BlockLinks_mbr;
fextl::pmr::named_monotonic_page_buffer_resource BlockLinks_mbr;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, FEXCore::Context::BlockDelinkerFunc>;
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
struct BlockEntry {
uint64_t HostCode;
fextl::vector<uint64_t> CodePages;
};
fextl::robin_map<uint64_t, BlockEntry> BlockList;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap();
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode, const LockToken&) {
const BlockEntry& AddBlockMapping(uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode, const LookupCacheWriteLockToken&) {
// This may replace an existing mapping
// NOTE: Generally no previous entry should exist, however there is one exception:
// If the backend updates the active thread's CodeBuffer, the new associated LookupCache
// may already contain the block address. Since is comparatively rare, we'll just leak
// one of the two blocks in this case.
BlockList[Address] = (uintptr_t)HostCode;
return BlockList.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, CodePages}).first->second;
}
std::optional<uintptr_t> FindBlock(uint64_t Address, const LockToken&) {
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
auto HostCode = BlockList.find(Address);
if (HostCode == BlockList.end()) {
return std::nullopt;
return nullptr;
}
return HostCode->second;
return &HostCode->second;
}
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LockToken&) {
bool Erase(uint64_t Address, const LookupCacheWriteLockToken&) {
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
it->second(it->first.HostLink);
}
// Remove from BlockList
return BlockList.erase(Address) != 0;
}
void InvalidateRange(uint64_t Start, uint64_t Length) {
auto lk = AcquireWriteLock();
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
Erase(Entry, lk);
}
}
CodePages.erase(lower, upper);
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
const FEXCore::Context::BlockDelinkerFunc& delinker, const LockToken&) {
const FEXCore::Context::BlockDelinkerFunc& delinker, const LookupCacheWriteLockToken&) {
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
}
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length, const LockToken&) {
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length, const LookupCacheWriteLockToken&) {
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length - 1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
@@ -107,7 +153,7 @@ struct GuestToHostMap {
return rv;
}
void ClearCache(const LockToken&);
void ClearCache(const LookupCacheWriteLockToken&);
};
class LookupCache {
@@ -122,122 +168,199 @@ public:
// Swaps out the underlying GuestToHostMap and clears all associated caches.
// This interface requires the previous CodeBuffer to be provided despite not using it. This ensures the shared write lock is still valid.
void ChangeGuestToHostMapping([[maybe_unused]] CPU::CodeBuffer& Prev, GuestToHostMap& NewMap) {
ClearThreadLocalCaches();
void ChangeGuestToHostMapping([[maybe_unused]] CPU::CodeBuffer& Prev, GuestToHostMap& NewMap, const LookupCacheWriteLockToken& lk) {
ClearThreadLocalCaches(lk);
Shared = &NewMap;
}
uintptr_t FindBlock(uint64_t Address) {
uintptr_t FindBlock(FEXCore::Core::InternalThreadState* Thread, uint64_t Address) {
// Try L1, no lock needed
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
if (L1Entry.GuestCode == Address) {
return L1Entry.HostCode;
}
// L2 and L3 need to be locked
auto lk = Shared->AcquireLock();
uintptr_t HostPtr {};
{
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheReadLockTime : nullptr);
auto lk = Shared->AcquireReadLock();
LockTime.reset();
// Try L2
const auto PageIndex = (Address & (VirtualMemSize - 1)) >> 12;
const auto PageOffset = Address & (0x0FFF);
if (!DisableL2Cache()) {
// Try L2
const auto PageIndex = (Address & (VirtualMemSize - 1)) >> 12;
const auto PageOffset = Address & (0x0FFF);
const auto Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
auto LocalPagePointer = Pointers[PageIndex];
const auto Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
auto LocalPagePointer = Pointers[PageIndex];
// Do we a page pointer for this address?
if (LocalPagePointer) {
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
// Do we a page pointer for this address?
if (LocalPagePointer) {
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
if (BlockPointers[PageOffset].GuestCode == Address) {
L1Entry.GuestCode = Address;
L1Entry.HostCode = BlockPointers[PageOffset].HostCode;
return L1Entry.HostCode;
if (BlockPointers[PageOffset].GuestCode == Address) {
L1Entry.GuestCode = Address;
L1Entry.HostCode = BlockPointers[PageOffset].HostCode;
HostPtr = L1Entry.HostCode;
}
}
}
if (!HostPtr) {
// Try L3
auto Entry = Shared->FindBlock(Address, lk);
if (Entry) {
CacheBlockMapping(Address, *Entry, false, lk);
HostPtr = Entry->HostCode;
}
}
}
// Try L3
auto HostCode = Shared->FindBlock(Address, lk);
if (HostCode) {
CacheBlockMapping(Address, HostCode.value());
return HostCode.value();
if (HostPtr && DynamicL1Cache()) {
UpdateDynamicL1Stats(Thread);
}
// Failed to find
return 0;
FEXCORE_PROFILE_INSTANT_INCREMENT(Thread, AccumulatedCacheMissCount, 1);
return HostPtr;
}
void UpdateDynamicL1Stats(FEXCore::Core::InternalThreadState* Thread) {
// If host pointer was found in L2 or L3, then add it to the counter.
// Keeping track not L1 misses, but specifically L2/L3 hits.
++L2L3CacheHits;
const auto CurrentTime = std::chrono::system_clock::now();
const auto Period = CurrentTime - LastPeriod;
if (Period >= SamplePeriod) {
// If larger than the sample period then check if we need to increase L1 cache size.
const double AveragePerSecond = static_cast<double>(L2L3CacheHits) /
static_cast<double>(std::chrono::duration_cast<std::chrono::milliseconds>(Period).count()) * 1000.0;
if (AveragePerSecond >= DynamicL1CacheIncreaseCountHeuristic()) {
if (CurrentL1Entries < MAX_L1_ENTRIES) {
CurrentL1Entries <<= 1;
L1PointerMask = CurrentL1Entries - 1;
// Update the thread's L1 pointer mask to increase how much cache it uses.
// Since we're in C-code, this is safe to update here.
Thread->CurrentFrame->State.L1Mask = GetScaledL1PointerMask();
}
} else if (AveragePerSecond < DynamicL1CacheDecreaseCountHeuristic()) {
if (CurrentL1Entries > MIN_L1_ENTRIES) {
CurrentL1Entries >>= 1;
L1PointerMask = CurrentL1Entries - 1;
// Madvise the entries that we are dropping. Gives the memory back to the OS.
LookupCacheEntry* FirstZeroL1Entry = &reinterpret_cast<LookupCacheEntry*>(L1Pointer)[CurrentL1Entries];
size_t ZeroMemorySize = (MAX_L1_ENTRIES - CurrentL1Entries) * sizeof(LookupCacheEntry);
FEXCore::Allocator::VirtualDontNeed(FirstZeroL1Entry, ZeroMemorySize, false);
// Update the thread's L1 pointer mask to increase how much cache it uses.
// Since we're in C-code, this is safe to update here.
Thread->CurrentFrame->State.L1Mask = GetScaledL1PointerMask();
}
}
// Update Last period to start again.
LastPeriod = CurrentTime;
L2L3CacheHits = 0;
}
}
GuestToHostMap* Shared = nullptr;
// Appends a list of Block {Address} to CodePages [Start, Start + Length)
// Returns true if new pages are marked as containing code
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireLock();
bool AddBlockExecutableRange(FEXCore::Core::InternalThreadState* Thread, const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
return Shared->AddBlockExecutableRange(Addresses, Start, Length, lk);
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode) {
auto lk = Shared->AcquireLock();
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
Shared->AddBlockMapping(Address, HostCode, lk);
const auto& Entry = Shared->AddBlockMapping(Address, CodePages, HostCode, lk);
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
CacheBlockMapping(Address, Entry, true, lk);
}
// NOTE: It's the caller's responsibility to call Erase() for all other
// GuestToHostMaps that share the same LookupCache. Otherwise, the
// L1/L2 caches will contain stale references to deallocated memory.
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address) {
auto lk = Shared->AcquireLock();
bool ErasedAny = Shared->Erase(Frame, Address, lk);
// Invalidates L1/L2 for a given guest block
void InvalidateCache(uint64_t Address, const LookupCacheWriteLockToken& lk) {
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
ErasedAny = true;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
}
// Do full map
Address = Address & (VirtualMemSize - 1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
if (!DisableL2Cache()) {
// Do full map
Address = Address & (VirtualMemSize - 1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
uintptr_t* Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return ErasedAny;
uintptr_t* Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return;
}
// Page exists, just set the offset to zero
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
}
// Page exists, just set the offset to zero
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
return true;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink, const FEXCore::Context::BlockDelinkerFunc& delinker) {
auto lk = Shared->AcquireLock();
// Invalidates all L1/L2 entries for all guest block that intersect the given range
bool InvalidateCacheRange(uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireWriteLock();
auto lower = CachedCodePages.lower_bound(Start >> 12);
auto upper = CachedCodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
InvalidateCache(Entry, lk);
}
}
bool ret = upper != lower;
CachedCodePages.erase(lower, upper);
return ret;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
const FEXCore::Context::BlockDelinkerFunc& delinker, const LookupCacheWriteLockToken& lk) {
Shared->AddBlockLink(GuestDestination, HostLink, delinker, lk);
}
void ClearCache();
void ClearL2Cache();
void ClearThreadLocalCaches();
void ClearCache(const LookupCacheWriteLockToken&);
void ClearL2Cache(const LookupCacheBaseLockToken&);
void ClearThreadLocalCaches(const LookupCacheWriteLockToken&);
uintptr_t GetL1Pointer() const {
return L1Pointer;
}
uintptr_t GetScaledL1PointerMask() const {
return L1PointerMask << FEXCore::ilog2(sizeof(LookupCache::LookupCacheEntry));
}
uintptr_t GetPagePointer() const {
return PagePointer;
}
@@ -245,9 +368,6 @@ public:
return VirtualMemSize;
}
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
// This needs to be taken before reads or writes to L2, L3, CodePages,
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// may only happen during cross thread invalidation (::Erase).
@@ -255,45 +375,52 @@ public:
// Some care is taken so that L1 lookups can be done without locks, and even tearing is unlikely to lead to a crash.
// This approach has not been fully vetted yet.
// Also note that L1 lookups might be inlined in the JIT Dispatcher and/or block ends.
auto AcquireLock() {
return Shared->AcquireLock();
auto AcquireWriteLock() {
return Shared->AcquireWriteLock();
}
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
L1Entry.HostCode = HostCode;
// Do ful map
auto FullAddress = Address;
Address = Address & (VirtualMemSize - 1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
uintptr_t* Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// We don't have a page pointer for this address
// Allocate one now if we can
uintptr_t NewPageBacking = AllocateBackingForPage();
if (!NewPageBacking) {
// Couldn't allocate, clear L2 and retry
ClearL2Cache();
CacheBlockMapping(Address, HostCode);
return;
}
Pointers[Address] = NewPageBacking;
LocalPagePointer = NewPageBacking;
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheBaseLockToken& lk) {
for (const auto& CodePage : Entry.CodePages) {
CachedCodePages[CodePage >> 12].insert(Address);
}
// Add the new pointer to the page block
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = Entry.HostCode;
// This silently replaces existing mappings
BlockPointers[PageOffset].GuestCode = FullAddress;
BlockPointers[PageOffset].HostCode = HostCode;
if (!DisableL2Cache() && !L1Only) {
// Do ful map
auto FullAddress = Address;
Address = Address & (VirtualMemSize - 1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
uintptr_t* Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// We don't have a page pointer for this address
// Allocate one now if we can
uintptr_t NewPageBacking = AllocateBackingForPage();
if (!NewPageBacking) {
// Couldn't allocate, clear L2 and retry
ClearL2Cache(lk);
CacheBlockMapping(Address, Entry, false, lk);
return;
}
Pointers[Address] = NewPageBacking;
LocalPagePointer = NewPageBacking;
}
// Add the new pointer to the page block
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
// This silently replaces existing mappings
BlockPointers[PageOffset].GuestCode = FullAddress;
BlockPointers[PageOffset].HostCode = Entry.HostCode;
}
}
uintptr_t AllocateBackingForPage() {
@@ -310,19 +437,38 @@ private:
return PageMemory + NewBase;
}
// Maps from a page index to all blocks in the page that have at some point been fetched into L1/L2
fextl::map<uint64_t, fextl::robin_set<uint64_t>> CachedCodePages;
uintptr_t PagePointer;
uintptr_t PageMemory;
uintptr_t L1Pointer;
uintptr_t L1PointerMask;
size_t TotalCacheSize;
// Start with 8k entries in L1 to give 128KB of L1 cache to each thread.
// Max out at 1 million entries to give each thread 16MB of L1 cache maximum.
constexpr static size_t MIN_L1_ENTRIES = 8 * 1024; // Must be a power of 2
constexpr static size_t MAX_L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t CODE_SIZE = 128 * 1024 * 1024;
constexpr static size_t SIZE_PER_PAGE = 4096 * sizeof(LookupCacheEntry);
constexpr static size_t L1_SIZE = L1_ENTRIES * sizeof(LookupCacheEntry);
constexpr static size_t SIZE_PER_PAGE = FEXCore::Utils::FEX_PAGE_SIZE * sizeof(LookupCacheEntry);
constexpr static size_t MAX_L1_SIZE = MAX_L1_ENTRIES * sizeof(LookupCacheEntry);
size_t AllocateOffset {};
FEXCore::Context::ContextImpl* ctx;
uint64_t VirtualMemSize {};
size_t CurrentL1Entries = MIN_L1_ENTRIES;
uint64_t L2L3CacheHits {};
std::chrono::time_point<std::chrono::system_clock> LastPeriod {};
constexpr static std::chrono::seconds SamplePeriod {1};
FEX_CONFIG_OPT(DynamicL1CacheIncreaseCountHeuristic, DYNAMICL1CACHEINCREASECOUNTHEURISTIC);
FEX_CONFIG_OPT(DynamicL1CacheDecreaseCountHeuristic, DYNAMICL1CACHEDECREASECOUNTHEURISTIC);
FEX_CONFIG_OPT(DynamicL1Cache, DYNAMICL1CACHE);
FEX_CONFIG_OPT(DisableL2Cache, DISABLEL2CACHE);
};
} // namespace FEXCore
File diff suppressed because it is too large. Load diff
+193 -117
View File
@@ -139,27 +139,28 @@ public:
FlushRegisterCache();
return _Jump(_TargetBlock);
}
IRPair<IROp_CondJump> CondJump(Ref _Cmp1, Ref _Cmp2, Ref _TrueBlock, Ref _FalseBlock, CondClassType _Cond = {COND_NEQ},
IRPair<IROp_CondJump> CondJump(Ref _Cmp1, Ref _Cmp2, Ref _TrueBlock, Ref _FalseBlock, CondClass _Cond = CondClass::NEQ,
IR::OpSize _CompareSize = OpSize::iInvalid) {
FlushRegisterCache();
return _CondJump(_Cmp1, _Cmp2, _TrueBlock, _FalseBlock, _Cond, _CompareSize);
}
IRPair<IROp_CondJump> CondJump(Ref ssa0, CondClassType cond = {COND_NEQ}) {
IRPair<IROp_CondJump> CondJump(Ref ssa0, CondClass cond = CondClass::NEQ) {
FlushRegisterCache();
return _CondJump(ssa0, cond);
}
IRPair<IROp_CondJump> CondJump(Ref ssa0, Ref ssa1, Ref ssa2, CondClassType cond = {COND_NEQ}) {
IRPair<IROp_CondJump> CondJump(Ref ssa0, Ref ssa1, Ref ssa2, CondClass cond = CondClass::NEQ) {
FlushRegisterCache();
return _CondJump(ssa0, ssa1, ssa2, cond);
}
IRPair<IROp_CondJump> CondJumpNZCV(CondClassType Cond) {
IRPair<IROp_CondJump> CondJumpNZCV(CondClass Cond) {
FlushRegisterCache();
return _CondJump(InvalidNode, InvalidNode, InvalidNode, InvalidNode, Cond, OpSize::iInvalid, true);
}
IRPair<IROp_CondJump> CondJumpBit(Ref Src, unsigned Bit, bool Set) {
FlushRegisterCache();
auto InlineConst = _InlineConstant(Bit);
return _CondJump(Src, InlineConst, InvalidNode, InvalidNode, {Set ? COND_TSTNZ : COND_TSTZ}, OpSize::iInvalid, false);
auto Cond = Set ? CondClass::TSTNZ : CondClass::TSTZ;
return _CondJump(Src, InlineConst, InvalidNode, InvalidNode, Cond, OpSize::iInvalid, false);
}
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP, BranchHint Hint = BranchHint::None) {
FlushRegisterCache();
@@ -251,7 +252,7 @@ public:
auto ExitBlock = CreateNewCodeBlockAfter(BackwardBlock);
auto DF = GetRFLAG(X86State::RFLAG_DF_RAW_LOC);
CondJump(DF, Zero, ForwardBlock, BackwardBlock, {COND_EQ});
CondJump(DF, Zero, ForwardBlock, BackwardBlock, CondClass::EQ);
for (auto D = 0; D < 2; ++D) {
SetCurrentCodeBlock(D ? BackwardBlock : ForwardBlock);
@@ -564,7 +565,7 @@ public:
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVXInsertScalar_CVT_Float_To_Float(OpcodeArgs);
RoundType TranslateRoundType(uint8_t Mode);
RoundMode TranslateRoundType(uint8_t Mode);
template<IR::OpSize ElementSize>
void InsertScalarRound(OpcodeArgs);
@@ -758,7 +759,6 @@ public:
void X87FXTRACT(OpcodeArgs);
void X87FYL2X(OpcodeArgs, bool IsFYL2XP1);
void X87LDENV(OpcodeArgs);
void X87LDSW(OpcodeArgs);
void X87ModifySTP(OpcodeArgs, bool Inc);
void X87OpHelper(OpcodeArgs, FEXCore::IR::IROps IROp, bool ZeroC2);
@@ -907,6 +907,10 @@ public:
void VPCLMULQDQOp(OpcodeArgs);
void CRC32(OpcodeArgs);
void Extrq_imm(OpcodeArgs);
void Insertq_imm(OpcodeArgs);
void Extrq(OpcodeArgs);
void Insertq(OpcodeArgs);
void BreakOp(OpcodeArgs, FEXCore::IR::BreakDefinition BreakDefinition);
void UnimplementedOp(OpcodeArgs);
@@ -985,7 +989,6 @@ public:
void AVX128_VPSIGN(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_UCOMISx(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VectorScalarInsertALU(OpcodeArgs, FEXCore::IR::IROps IROp, IR::OpSize ElementSize);
Ref AVX128_VFCMPImpl(IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t CompType);
void AVX128_VFCMP(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_InsertScalarFCMP(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_MOVBetweenGPR_FPR(OpcodeArgs);
@@ -1004,9 +1007,7 @@ public:
void AVX128_VINSERT(OpcodeArgs);
void AVX128_VINSERTPS(OpcodeArgs);
Ref AVX128_PHSUBImpl(Ref Src1, Ref Src2, size_t ElementSize);
void AVX128_VPHSUB(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPHSUBSW(OpcodeArgs);
void AVX128_VADDSUBP(OpcodeArgs, IR::OpSize ElementSize);
@@ -1098,8 +1099,8 @@ public:
void AVX128_VFMAScalarImpl(OpcodeArgs, IROps IROp, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx);
void AVX128_VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx);
RefPair AVX128_VPGatherQPSImpl(Ref Dest, Ref Mask, RefVSIB VSIB);
RefPair AVX128_VPGatherImpl(OpSize Size, OpSize ElementLoadSize, OpSize AddrElementSize, RefPair Dest, RefPair Mask, RefVSIB VSIB);
RefPair AVX128_VPGatherQPSImpl(OpcodeArgs, Ref Dest, Ref Mask, RefVSIB VSIB);
RefPair AVX128_VPGatherImpl(OpcodeArgs, OpSize Size, OpSize ElementLoadSize, OpSize AddrElementSize, RefPair Dest, RefPair Mask, RefVSIB VSIB);
void AVX128_VPGATHER(OpcodeArgs, OpSize AddrElementSize);
@@ -1109,8 +1110,8 @@ public:
// End of AVX 128-bit implementation
// AVX 256-bit operations
void StoreResult_WithAVXInsert(VectorOpType Type, FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, Ref Value,
IR::OpSize Align, MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
void StoreResult_WithAVXInsert(VectorOpType Type, RegClass Class, FEXCore::X86Tables::DecodedOp Op, Ref Value,
IR::OpSize Align = IR::OpSize::iInvalid, MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
if (Op->Dest.IsGPR() && Op->Dest.Data.GPR.GPR >= X86State::REG_XMM_0 && Op->Dest.Data.GPR.GPR <= X86State::REG_XMM_15 &&
GetGuestVectorLength() == OpSize::i256Bit && Type == VectorOpType::SSE) {
const auto gpr = Op->Dest.Data.GPR.GPR;
@@ -1161,7 +1162,7 @@ public:
}
}
void StoreContextHelper(IR::OpSize Size, RegisterClassType Class, Ref Value, uint32_t Offset) {
void StoreContextHelper(IR::OpSize Size, RegClass Class, Ref Value, uint32_t Offset) {
// For i128Bit, we won't see a normal Constant to inline, but as a special
// case we can replace with a 2x64-bit store which can use inline zeroes.
if (Size == OpSize::i128Bit) {
@@ -1173,7 +1174,7 @@ public:
if (Const->Constant == IR::NamedVectorConstant::NAMED_VECTOR_ZERO) {
Ref Zero = _Constant(0);
Ref STP = _StoreContextPair(IR::OpSize::i64Bit, GPRClass, Zero, Zero, Offset);
Ref STP = _StoreContextPair(IR::OpSize::i64Bit, RegClass::GPR, Zero, Zero, Offset);
// XXX: This works around InlineConstant not having an associated
// register class, else we'd just do InlineConstant above.
@@ -1229,16 +1230,16 @@ public:
if (Index >= GPR0Index && Index <= GPR15Index) {
Ref R = _StoreRegister(Value, GPRSize);
R->Reg = PhysicalRegister(GPRFixedClass, Index - GPR0Index).Raw;
R->Reg = PhysicalRegister(RegClass::GPRFixed, Index - GPR0Index).Raw;
} else if (Index == PFIndex) {
_StorePF(Value, GPRSize);
} else if (Index == AFIndex) {
_StoreAF(Value, GPRSize);
} else if (Index >= FPR0Index && Index <= FPR15Index) {
Ref R = _StoreRegister(Value, VectorSize);
R->Reg = PhysicalRegister(FPRFixedClass, Index - FPR0Index).Raw;
R->Reg = PhysicalRegister(RegClass::FPRFixed, Index - FPR0Index).Raw;
} else if (Index == DFIndex) {
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(Core::CPUState, flags[X86State::RFLAG_DF_RAW_LOC]));
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(Core::CPUState, flags[X86State::RFLAG_DF_RAW_LOC]));
} else {
bool Partial = RegCache.Partial & (1ull << Index);
auto Size = Partial ? OpSize::i64Bit : CacheIndexToOpSize(Index);
@@ -1263,7 +1264,7 @@ public:
StoreContextHelper(Size, Class, Value, Offset);
// If Partial and MMX register, then we need to store all 1s in bits 64-80
if (Partial && Index >= MM0Index && Index <= MM7Index) {
_StoreContext(OpSize::i16Bit, IR::GPRClass, Constant(0xFFFF), Offset + 8);
_StoreContextGPR(OpSize::i16Bit, Constant(0xFFFF), Offset + 8);
}
}
}
@@ -1553,23 +1554,63 @@ private:
AddressMode DecodeAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, MemoryAccessType AccessType, bool IsLoad);
Ref LoadSource(RegisterClassType Class, const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags,
Ref LoadSource(RegClass Class, const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags,
const LoadSourceOptions& Options = {});
Ref LoadSource_WithOpSize(RegisterClassType Class, const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand,
IR::OpSize OpSize, uint32_t Flags, const LoadSourceOptions& Options = {});
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op,
const FEXCore::X86Tables::DecodedOperand& Operand, const Ref Src, IR::OpSize OpSize, IR::OpSize Align,
MemoryAccessType AccessType = MemoryAccessType::DEFAULT);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, const FEXCore::X86Tables::DecodedOperand& Operand,
const Ref Src, IR::OpSize Align, MemoryAccessType AccessType = MemoryAccessType::DEFAULT);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, const Ref Src, IR::OpSize Align,
Ref LoadSourceGPR(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags,
const LoadSourceOptions& Options = {}) {
return LoadSource(RegClass::GPR, Op, Operand, Flags, Options);
}
Ref LoadSourceFPR(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags,
const LoadSourceOptions& Options = {}) {
return LoadSource(RegClass::FPR, Op, Operand, Flags, Options);
}
Ref LoadSource_WithOpSize(RegClass Class, const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, IR::OpSize OpSize,
uint32_t Flags, const LoadSourceOptions& Options = {});
Ref LoadSourceGPR_WithOpSize(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, IR::OpSize OpSize, uint32_t Flags,
const LoadSourceOptions& Options = {}) {
return LoadSource_WithOpSize(RegClass::GPR, Op, Operand, OpSize, Flags, Options);
}
Ref LoadSourceFPR_WithOpSize(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, IR::OpSize OpSize, uint32_t Flags,
const LoadSourceOptions& Options = {}) {
return LoadSource_WithOpSize(RegClass::FPR, Op, Operand, OpSize, Flags, Options);
}
void StoreResult_WithOpSize(RegClass Class, X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src, IR::OpSize OpSize,
IR::OpSize Align, MemoryAccessType AccessType = MemoryAccessType::DEFAULT);
void StoreResultGPR_WithOpSize(X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src, IR::OpSize OpSize,
IR::OpSize Align = IR::OpSize::iInvalid, MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
StoreResult_WithOpSize(RegClass::GPR, Op, Operand, Src, OpSize, Align, AccessType);
}
void StoreResultFPR_WithOpSize(X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src, IR::OpSize OpSize,
IR::OpSize Align = IR::OpSize::iInvalid, MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
StoreResult_WithOpSize(RegClass::FPR, Op, Operand, Src, OpSize, Align, AccessType);
}
void StoreResult(RegClass Class, X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src, OpSize Align,
MemoryAccessType AccessType = MemoryAccessType::DEFAULT);
void StoreResultGPR(X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src, OpSize Align = OpSize::iInvalid,
MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
StoreResult(RegClass::GPR, Op, Operand, Src, Align, AccessType);
}
void StoreResultFPR(X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src, OpSize Align = OpSize::iInvalid,
MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
StoreResult(RegClass::FPR, Op, Operand, Src, Align, AccessType);
}
void StoreResult(RegClass Class, X86Tables::DecodedOp Op, Ref Src, OpSize Align, MemoryAccessType AccessType = MemoryAccessType::DEFAULT);
void StoreResultGPR(X86Tables::DecodedOp Op, Ref Src, OpSize Align = OpSize::iInvalid, MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
StoreResult(RegClass::GPR, Op, Src, Align, AccessType);
}
void StoreResultFPR(X86Tables::DecodedOp Op, Ref Src, OpSize Align = OpSize::iInvalid, MemoryAccessType AccessType = MemoryAccessType::DEFAULT) {
StoreResult(RegClass::FPR, Op, Src, Align, AccessType);
}
// In several instances, it's desirable to get a base address with the segment offset
// applied to it. This pulls all the common-case appending into a single set of functions.
[[nodiscard]]
Ref MakeSegmentAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, IR::OpSize OpSize) {
Ref Mem = LoadSource_WithOpSize(GPRClass, Op, Operand, OpSize, Op->Flags, {.LoadData = false});
Ref Mem = LoadSourceGPR_WithOpSize(Op, Operand, OpSize, Op->Flags, {.LoadData = false});
return AppendSegmentOffset(Mem, Op->Flags);
}
[[nodiscard]]
@@ -1614,6 +1655,9 @@ private:
return IR::SizeToOpSize(GetSrcSize(Op));
}
[[nodiscard]]
IR::OpSize GetStringOpSize(X86Tables::DecodedOp Op) const;
// Set flag tracking to prepare for an operation that directly writes NZCV.
void HandleNZCVWrite() {
CachedNZCV = nullptr;
@@ -1805,14 +1849,15 @@ private:
// For DF, we need to transform 0/1 into 1/-1
StoreDF(_SubShift(OpSize::i64Bit, Constant(1), Value, ShiftType::LSL, 1));
} else if (BitOffset == FEXCore::X86State::RFLAG_TF_RAW_LOC) {
auto PackedTF = _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
// An exception should still be raised after an instruction that unsets TF, leave the unblocked bit set but unset
// the TF bit to cause such behaviour. The handling code at the start of the next block will then unset the
// unblocked bit before raising the exception.
auto NewPackedTF = _Select(FEXCore::IR::COND_EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContext(OpSize::i8Bit, GPRClass, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
auto NewPackedTF =
_Select(OpSize::i64Bit, OpSize::i64Bit, CondClass::EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
} else {
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
}
}
@@ -1838,12 +1883,12 @@ private:
}
[[nodiscard]]
static CondClassType CondForNZCVBit(unsigned BitOffset, bool Invert) {
static CondClass CondForNZCVBit(unsigned BitOffset, bool Invert) {
switch (BitOffset) {
case X86State::RFLAG_SF_RAW_LOC: return {Invert ? COND_PL : COND_MI};
case X86State::RFLAG_ZF_RAW_LOC: return {Invert ? COND_NEQ : COND_EQ};
case X86State::RFLAG_CF_RAW_LOC: return {Invert ? COND_ULT : COND_UGE};
case X86State::RFLAG_OF_RAW_LOC: return {Invert ? COND_FNU : COND_FU};
case X86State::RFLAG_SF_RAW_LOC: return Invert ? CondClass::PL : CondClass::MI;
case X86State::RFLAG_ZF_RAW_LOC: return Invert ? CondClass::NEQ : CondClass::EQ;
case X86State::RFLAG_CF_RAW_LOC: return Invert ? CondClass::ULT : CondClass::UGE;
case X86State::RFLAG_OF_RAW_LOC: return Invert ? CondClass::FNU : CondClass::FU;
default: FEX_UNREACHABLE;
}
}
@@ -1874,11 +1919,11 @@ private:
}
[[nodiscard]]
static RegisterClassType CacheIndexClass(int Index) {
static RegClass CacheIndexClass(int Index) {
if ((Index >= MM0Index && Index <= MM7Index) || Index >= FPR0Index) {
return FPRClass;
return RegClass::FPR;
} else {
return GPRClass;
return RegClass::GPR;
}
}
@@ -1910,14 +1955,14 @@ private:
RegCache.Written &= ~Bit;
}
Ref LoadRegCache(uint64_t Offset, uint8_t Index, RegisterClassType RegClass, IR::OpSize Size) {
Ref LoadRegCache(uint64_t Offset, uint8_t Index, RegClass Class, IR::OpSize Size) {
LOGMAN_THROW_A_FMT(Index < 64, "valid index");
uint64_t Bit = (1ull << (uint64_t)Index);
if (Size == OpSize::i128Bit && (RegCache.Partial & Bit)) {
// We need to load the full register extend if we previously did a partial access.
Ref Value = RegCache.Value[Index];
Ref Full = _LoadContext(Size, RegClass, Offset);
Ref Full = _LoadContext(Size, Class, Offset);
// If we did a partial store, we're inserting into the full register
if (RegCache.Written & Bit) {
@@ -1931,7 +1976,7 @@ private:
if (Index == DFIndex) {
RegCache.Value[Index] = _LoadDF();
} else if ((Index >= MM0Index && Index <= MM7Index) || Index >= AVXHigh0Index) {
RegCache.Value[Index] = _LoadContext(Size, RegClass, Offset);
RegCache.Value[Index] = _LoadContext(Size, Class, Offset);
// We may have done a partial load, this requires special handling.
if (Size == OpSize::i64Bit) {
@@ -1942,7 +1987,7 @@ private:
} else if (Index == AFIndex) {
RegCache.Value[Index] = _LoadAF(Size);
} else {
RegCache.Value[Index] = _LoadRegister(Offset, RegClass, Size);
RegCache.Value[Index] = _LoadRegister(Offset, Class, Size);
}
RegCache.Cached |= Bit;
@@ -1951,21 +1996,21 @@ private:
return RegCache.Value[Index];
}
RefPair AllocatePair(FEXCore::IR::RegisterClassType Class, IR::OpSize Size) {
if (Class == FPRClass) {
RefPair AllocatePair(RegClass Class, IR::OpSize Size) {
if (Class == RegClass::FPR) {
return {_AllocateFPR(Size, Size), _AllocateFPR(Size, Size)};
} else {
return {_AllocateGPR(false), _AllocateGPR(false)};
}
}
RefPair LoadContextPair_Uncached(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, unsigned Offset) {
RefPair LoadContextPair_Uncached(RegClass Class, IR::OpSize Size, unsigned Offset) {
RefPair Values = AllocatePair(Class, Size);
_LoadContextPair(Size, Class, Offset, Values.Low, Values.High);
return Values;
}
RefPair LoadRegCachePair(uint64_t Offset, uint8_t Index, RegisterClassType RegClass, IR::OpSize Size) {
RefPair LoadRegCachePair(uint64_t Offset, uint8_t Index, RegClass Class, IR::OpSize Size) {
LOGMAN_THROW_A_FMT(Index != DFIndex, "must be pairable");
LOGMAN_THROW_A_FMT(Size != IR::OpSize::iUnsized, "Invalid size!");
@@ -1973,7 +2018,7 @@ private:
uint64_t Bits = (3ull << (uint64_t)Index);
const auto SizeInt = IR::OpSizeToSize(Size);
if (((RegCache.Partial | RegCache.Cached) & Bits) == 0 && ((Offset / SizeInt) < 64)) {
auto Values = LoadContextPair_Uncached(RegClass, Size, Offset);
auto Values = LoadContextPair_Uncached(Class, Size, Offset);
RegCache.Value[Index] = Values.Low;
RegCache.Value[Index + 1] = Values.High;
RegCache.Cached |= Bits;
@@ -1985,13 +2030,13 @@ private:
// Fallback on a pair of loads
return {
.Low = LoadRegCache(Offset, Index, RegClass, Size),
.High = LoadRegCache(Offset + SizeInt, Index + 1, RegClass, Size),
.Low = LoadRegCache(Offset, Index, Class, Size),
.High = LoadRegCache(Offset + SizeInt, Index + 1, Class, Size),
};
}
Ref LoadGPR(uint8_t Reg) {
return LoadRegCache(Reg, GPR0Index + Reg, GPRClass, GetGPROpSize());
return LoadRegCache(Reg, GPR0Index + Reg, RegClass::GPR, GetGPROpSize());
}
Ref LoadContext(IR::OpSize Size, uint8_t Index) {
@@ -2007,7 +2052,7 @@ private:
}
Ref LoadXMMRegister(uint8_t Reg) {
return LoadRegCache(Reg, FPR0Index + Reg, FPRClass, GetGuestVectorLength());
return LoadRegCache(Reg, FPR0Index + Reg, RegClass::FPR, GetGuestVectorLength());
}
Ref LoadDF() {
@@ -2062,7 +2107,7 @@ private:
// Recover the sign bit, it is the logical DF value
return _Lshr(OpSize::i64Bit, LoadDF(), Constant(63));
} else {
return _LoadContext(OpSize::i8Bit, GPRClass, offsetof(Core::CPUState, flags[BitOffset]));
return _LoadContextGPR(OpSize::i8Bit, offsetof(Core::CPUState, flags[BitOffset]));
}
}
@@ -2079,18 +2124,18 @@ private:
}
// Safe version of NZCVSelect that handles inverted carries automatically.
Ref NZCVSelect(OpSize OpSize, CondClassType Cond, Ref TrueV, Ref FalseV, bool CarryIsInverted = false) {
Ref NZCVSelect(OpSize OpSize, CondClass Cond, Ref TrueV, Ref FalseV, bool CarryIsInverted = false) {
switch (Cond) {
case IR::COND_UGE: /* cs */
case IR::COND_ULT: /* cc */
case CondClass::UGE: /* cs */
case CondClass::ULT: /* cc */
// Invert the condition to match our expectations.
if (CarryIsInverted != CFInverted) {
Cond = {Cond == COND_UGE ? COND_ULT : COND_UGE};
Cond = (Cond == CondClass::UGE) ? CondClass::ULT : CondClass::UGE;
}
break;
case IR::COND_UGT: /* hi */
case IR::COND_ULE: /* ls */
case CondClass::UGT: /* hi */
case CondClass::ULE: /* ls */
// No clever optimization we can do here, rectify carry itself.
RectifyCarryInvert(CarryIsInverted);
break;
@@ -2179,7 +2224,7 @@ private:
HandleNZCV_RMW();
CalculatePF(_ShiftFlags(OpSizeFromSrc(Op), Result, Dest, Shift, Src, OldPF, CFInverted));
StoreResult(GPRClass, Op, Result, OpSize::iInvalid);
StoreResultGPR(Op, Result);
}
// Helper to derive Dest by a given builder-using Expression with the opcode
@@ -2252,8 +2297,7 @@ private:
CachedIndexedNamedVectorConstants.clear();
}
std::optional<CondClassType> DecodeNZCVCondition(uint8_t OP);
Ref SelectBit(Ref Cmp, IR::OpSize ResultSize, Ref TrueValue, Ref FalseValue);
std::optional<CondClass> DecodeNZCVCondition(uint8_t OP);
Ref SelectCC0All1(uint8_t OP);
/**
@@ -2267,8 +2311,8 @@ private:
if (Size != OpSize::i32Bit) {
return;
}
auto Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags);
StoreResult(GPRClass, Op, Dest, OpSize::iInvalid);
auto Dest = LoadSourceGPR(Op, Op->Dest, Op->Flags);
StoreResultGPR(Op, Dest);
}
using ZeroShiftFunctionPtr = void (OpDispatchBuilder::*)(FEXCore::X86Tables::DecodedOp Op);
@@ -2303,7 +2347,7 @@ private:
///< Jump to zeroshift block or end block depending on if it was provided.
IRPair<IROp_CodeBlock> TailHandling = ZeroShiftResult ? ZeroShiftBlock : EndBlock;
CondJump(Shift, Zero, TailHandling, SetBlock, {COND_EQ});
CondJump(Shift, Zero, TailHandling, SetBlock, CondClass::EQ);
SetCurrentCodeBlock(SetBlock);
StartNewBlock();
@@ -2346,9 +2390,7 @@ private:
void CalculateFlags_MUL(IR::OpSize SrcSize, Ref Res, Ref High);
void CalculateFlags_UMUL(Ref High);
void CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res);
void CalculateFlags_ShiftLeft(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2);
void CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
void CalculateFlags_ShiftRight(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2);
void CalculateFlags_ShiftRightImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
void CalculateFlags_ShiftRightDoubleImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
void CalculateFlags_ShiftRightImmediateCommon(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
@@ -2365,7 +2407,7 @@ private:
LOGMAN_THROW_A_FMT(MMXState == MMXState_X87, "Expected state to be x87");
_StackForceSlow();
SetX87Top(Constant(0)); // top reset to zero
_StoreContext(OpSize::i8Bit, GPRClass, Constant(0xFFFFUL), offsetof(FEXCore::Core::CPUState, AbridgedFTW));
_StoreContextGPR(OpSize::i8Bit, Constant(0xFFFFUL), offsetof(FEXCore::Core::CPUState, AbridgedFTW));
MMXState = MMXState_MMX;
}
@@ -2402,44 +2444,62 @@ private:
IROp_IRHeader* CurrentHeader {};
[[nodiscard]]
bool IsTSOEnabled(FEXCore::IR::RegisterClassType Class) const {
bool IsTSOEnabled(RegClass Class) const {
if (ForceTSO == ForceTSOMode::ForceEnabled) {
return true;
} else if (ForceTSO == ForceTSOMode::ForceDisabled) {
return false;
} else if (Class == FPRClass) {
} else if (Class == RegClass::FPR) {
return CTX->IsVectorAtomicTSOEnabled();
} else {
return CTX->IsAtomicTSOEnabled();
}
}
Ref _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref Addr, Ref Value, IR::OpSize Align = IR::OpSize::i8Bit) {
Ref _StoreMemAutoTSO(RegClass Class, OpSize Size, Ref Addr, Ref Value, OpSize Align = OpSize::i8Bit) {
if (IsTSOEnabled(Class)) {
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MemOffsetType::SXTX, 1);
} else {
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MemOffsetType::SXTX, 1);
}
}
Ref _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref ssa0, IR::OpSize Align = IR::OpSize::i8Bit) {
if (IsTSOEnabled(Class)) {
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
} else {
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
Ref _StoreMemGPRAutoTSO(OpSize Size, Ref Addr, Ref Value, OpSize Align = OpSize::i8Bit) {
return _StoreMemAutoTSO(RegClass::GPR, Size, Addr, Value, Align);
}
Ref _StoreMemFPRAutoTSO(OpSize Size, Ref Addr, Ref Value, OpSize Align = OpSize::i8Bit) {
return _StoreMemAutoTSO(RegClass::FPR, Size, Addr, Value, Align);
}
Ref _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, IR::OpSize Align = IR::OpSize::i8Bit) {
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
Ref _LoadMemAutoTSO(RegClass Class, OpSize Size, Ref ssa0, OpSize Align = OpSize::i8Bit) {
if (IsTSOEnabled(Class)) {
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MemOffsetType::SXTX, 1);
} else {
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MemOffsetType::SXTX, 1);
}
}
Ref _LoadMemGPRAutoTSO(OpSize Size, Ref ssa0, OpSize Align = OpSize::i8Bit) {
return _LoadMemAutoTSO(RegClass::GPR, Size, ssa0, Align);
}
Ref _LoadMemFPRAutoTSO(OpSize Size, Ref ssa0, OpSize Align = OpSize::i8Bit) {
return _LoadMemAutoTSO(RegClass::FPR, Size, ssa0, Align);
}
Ref _LoadMemAutoTSO(RegClass Class, OpSize Size, const AddressMode& A, OpSize Align = OpSize::i8Bit) {
const bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
const auto B = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != RegClass::GPR, Size);
if (AtomicTSO) {
return _LoadMemTSO(Class, Size, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
return _LoadMemTSO(Class, Size, B.Base, B.Index, Align, B.IndexType, B.IndexScale);
} else {
return _LoadMem(Class, Size, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
return _LoadMem(Class, Size, B.Base, B.Index, Align, B.IndexType, B.IndexScale);
}
}
Ref _LoadMemGPRAutoTSO(OpSize Size, const AddressMode& A, OpSize Align = OpSize::i8Bit) {
return _LoadMemAutoTSO(RegClass::GPR, Size, A, Align);
}
Ref _LoadMemFPRAutoTSO(OpSize Size, const AddressMode& A, OpSize Align = OpSize::i8Bit) {
return _LoadMemAutoTSO(RegClass::FPR, Size, A, Align);
}
AddressMode SelectPairAddressMode(AddressMode A, IR::OpSize Size) {
LOGMAN_THROW_A_FMT(Size != IR::OpSize::iUnsized, "Invalid size!");
@@ -2457,56 +2517,72 @@ private:
}
RefPair LoadMemPair(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref Base, unsigned Offset) {
RefPair LoadMemPair(RegClass Class, OpSize Size, Ref Base, uint32_t Offset) {
RefPair Values = AllocatePair(Class, Size);
_LoadMemPair(Class, Size, Base, Offset, Values.Low, Values.High);
return Values;
}
RefPair LoadMemPairFPR(OpSize Size, Ref Base, uint32_t Offset) {
return LoadMemPair(RegClass::FPR, Size, Base, Offset);
}
RefPair _LoadMemPairAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, IR::OpSize Align = IR::OpSize::i8Bit) {
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
RefPair _LoadMemPairAutoTSO(RegClass Class, OpSize Size, const AddressMode& A, OpSize Align = OpSize::i8Bit) {
const bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
// Use ldp if possible, otherwise fallback on two loads.
if (!AtomicTSO && !A.Segment && Size >= OpSize::i32Bit & Size <= OpSize::i128Bit) {
A = SelectPairAddressMode(A, Size);
return LoadMemPair(Class, Size, A.Base, A.Offset);
} else {
AddressMode HighA = A;
HighA.Offset += 16;
return {
.Low = _LoadMemAutoTSO(Class, Size, A, Align),
.High = _LoadMemAutoTSO(Class, Size, HighA, Align),
};
if (!AtomicTSO && !A.Segment && Size >= OpSize::i32Bit && Size <= OpSize::i128Bit) {
const auto B = SelectPairAddressMode(A, Size);
return LoadMemPair(Class, Size, B.Base, B.Offset);
}
AddressMode HighA = A;
HighA.Offset += 16;
return {
.Low = _LoadMemAutoTSO(Class, Size, A, Align),
.High = _LoadMemAutoTSO(Class, Size, HighA, Align),
};
}
RefPair _LoadMemPairFPRAutoTSO(OpSize Size, const AddressMode& A, OpSize Align = OpSize::i8Bit) {
return _LoadMemPairAutoTSO(RegClass::FPR, Size, A, Align);
}
Ref _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, Ref Value, IR::OpSize Align = IR::OpSize::i8Bit) {
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
Ref _StoreMemAutoTSO(RegClass Class, OpSize Size, const AddressMode& A, Ref Value, OpSize Align = OpSize::i8Bit) {
const bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
const auto B = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != RegClass::GPR, Size);
if (AtomicTSO) {
return _StoreMemTSO(Class, Size, Value, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
return _StoreMemTSO(Class, Size, Value, B.Base, B.Index, Align, B.IndexType, B.IndexScale);
} else {
return _StoreMem(Class, Size, Value, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
return _StoreMem(Class, Size, Value, B.Base, B.Index, Align, B.IndexType, B.IndexScale);
}
}
Ref _StoreMemGPRAutoTSO(OpSize Size, const AddressMode& A, Ref Value, OpSize Align = OpSize::i8Bit) {
return _StoreMemAutoTSO(RegClass::GPR, Size, A, Value, Align);
}
Ref _StoreMemFPRAutoTSO(OpSize Size, const AddressMode& A, Ref Value, OpSize Align = OpSize::i8Bit) {
return _StoreMemAutoTSO(RegClass::FPR, Size, A, Value, Align);
}
void _StoreMemPairAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, Ref Value1, Ref Value2,
IR::OpSize Align = IR::OpSize::i8Bit) {
void _StoreMemPairAutoTSO(RegClass Class, OpSize Size, const AddressMode& A, Ref Value1, Ref Value2, OpSize Align = OpSize::i8Bit) {
const auto SizeInt = IR::OpSizeToSize(Size);
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
const bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
// Use stp if possible, otherwise fallback on two stores.
if (!AtomicTSO && !A.Segment && Size >= OpSize::i32Bit & Size <= OpSize::i128Bit) {
A = SelectPairAddressMode(A, Size);
_StoreMemPair(Class, Size, Value1, Value2, A.Base, A.Offset);
if (!AtomicTSO && !A.Segment && Size >= OpSize::i32Bit && Size <= OpSize::i128Bit) {
const auto B = SelectPairAddressMode(A, Size);
_StoreMemPair(Class, Size, Value1, Value2, B.Base, B.Offset);
} else {
_StoreMemAutoTSO(Class, Size, A, Value1, OpSize::i8Bit);
A.Offset += SizeInt;
_StoreMemAutoTSO(Class, Size, A, Value2, OpSize::i8Bit);
auto B = A;
_StoreMemAutoTSO(Class, Size, B, Value1, OpSize::i8Bit);
B.Offset += SizeInt;
_StoreMemAutoTSO(Class, Size, B, Value2, OpSize::i8Bit);
}
}
void _StoreMemPairFPRAutoTSO(OpSize Size, const AddressMode& A, Ref Value1, Ref Value2, OpSize Align = OpSize::i8Bit) {
return _StoreMemPairAutoTSO(RegClass::FPR, Size, A, Value1, Value2, Align);
}
Ref Pop(IR::OpSize Size, Ref SP_RMW) {
Ref Value = _AllocateGPR(false);
@@ -35,20 +35,16 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_LoadSource_WithOpSize(
} else {
LOGMAN_THROW_A_FMT(IsOperandMem(Operand, true), "only memory sources");
AddressMode A = DecodeAddress(Op, Operand, AccessType, true /* IsLoad */);
AddressMode HighA = A;
HighA.Offset += 16;
if (Operand.IsSIB()) {
const bool IsVSIB = (Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0;
LOGMAN_THROW_A_FMT(!IsVSIB, "VSIB uses LoadVSIB instead");
}
const AddressMode A = DecodeAddress(Op, Operand, AccessType, true /* IsLoad */);
if (NeedsHigh) {
return _LoadMemPairAutoTSO(FPRClass, OpSize::i128Bit, A, OpSize::i8Bit);
return _LoadMemPairFPRAutoTSO(OpSize::i128Bit, A, OpSize::i8Bit);
} else {
return {.Low = _LoadMemAutoTSO(FPRClass, OpSize::i128Bit, A, OpSize::i8Bit)};
return {.Low = _LoadMemFPRAutoTSO(OpSize::i128Bit, A, OpSize::i8Bit)};
}
}
}
@@ -95,9 +91,9 @@ void OpDispatchBuilder::AVX128_StoreResult_WithOpSize(FEXCore::X86Tables::Decode
AddressMode A = DecodeAddress(Op, Operand, AccessType, false /* IsLoad */);
if (Src.High) {
_StoreMemPairAutoTSO(FPRClass, OpSize::i128Bit, A, Src.Low, Src.High, OpSize::i8Bit);
_StoreMemPairFPRAutoTSO(OpSize::i128Bit, A, Src.Low, Src.High, OpSize::i8Bit);
} else {
_StoreMemAutoTSO(FPRClass, OpSize::i128Bit, A, Src.Low, OpSize::i8Bit);
_StoreMemFPRAutoTSO(OpSize::i128Bit, A, Src.Low, OpSize::i8Bit);
}
}
}
@@ -151,13 +147,13 @@ void OpDispatchBuilder::AVX128_VMOVScalarImpl(OpcodeArgs, IR::OpSize ElementSize
AVX128_StoreResult_WithOpSize(Op, Op->Dest, RefPair {.Low = Result, .High = High});
} else if (Op->Dest.IsGPR()) {
// VMOVSS/SD xmm1, mem32/mem64
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], ElementSize, Op->Flags);
Ref Src = LoadSourceFPR_WithOpSize(Op, Op->Src[1], ElementSize, Op->Flags);
auto High = LoadZeroVector(OpSize::i128Bit);
AVX128_StoreResult_WithOpSize(Op, Op->Dest, RefPair {.Low = Src, .High = High});
} else {
// VMOVSS/SD mem32/mem64, xmm1
auto Src = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, false);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src.Low, ElementSize, OpSize::iInvalid);
StoreResultFPR_WithOpSize(Op, Op->Dest, Src.Low, ElementSize);
}
}
@@ -351,7 +347,7 @@ void OpDispatchBuilder::AVX128_MOVVectorNT(OpcodeArgs) {
if (Op->Dest.IsGPR()) {
///< MOVNTDQA load non-temporal comes from SSE4.1 and is extended by AVX/AVX2.
RefPair Src {};
Ref SrcAddr = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, {.LoadData = false});
Ref SrcAddr = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.LoadData = false});
Src.Low = _VLoadNonTemporal(OpSize::i128Bit, SrcAddr, 0);
if (Is128Bit) {
@@ -362,7 +358,7 @@ void OpDispatchBuilder::AVX128_MOVVectorNT(OpcodeArgs) {
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Src);
} else {
auto Src = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, !Is128Bit, MemoryAccessType::STREAM);
Ref Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
Ref Dest = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.LoadData = false});
if (Is128Bit) {
// Single store non-temporal for 128-bit operations.
@@ -379,7 +375,7 @@ void OpDispatchBuilder::AVX128_MOVQ(OpcodeArgs) {
if (Op->Src[0].IsGPR()) {
Src = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, false);
} else {
Src.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::i64Bit, Op->Flags);
Src.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::i64Bit, Op->Flags);
}
// This instruction is a bit special that if the destination is a register then it'll ZEXT the 64bit source to 256bit
@@ -390,7 +386,7 @@ void OpDispatchBuilder::AVX128_MOVQ(OpcodeArgs) {
Src.High = ZeroVector;
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Src);
} else {
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src.Low, OpSize::i64Bit, OpSize::i64Bit);
StoreResultFPR_WithOpSize(Op, Op->Dest, Src.Low, OpSize::i64Bit, OpSize::i64Bit);
}
}
@@ -399,7 +395,7 @@ void OpDispatchBuilder::AVX128_VMOVLP(OpcodeArgs) {
if (!Op->Dest.IsGPR()) {
///< VMOVLPS/PD mem64, xmm1
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src1.Low, OpSize::i64Bit, OpSize::i64Bit);
StoreResultFPR_WithOpSize(Op, Op->Dest, Src1.Low, OpSize::i64Bit, OpSize::i64Bit);
} else if (!Op->Src[1].IsGPR()) {
///< VMOVLPS/PD xmm1, xmm2, mem64
// Bits[63:0] come from Src2[63:0]
@@ -463,7 +459,7 @@ void OpDispatchBuilder::AVX128_VMOVDDUP(OpcodeArgs) {
// 128-bit operation only loads 8-bytes.
// 256-bit operation loads a full 32-bytes.
if (Is128Bit) {
Src.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::i64Bit, Op->Flags);
Src.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::i64Bit, Op->Flags);
} else {
Src = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, true);
}
@@ -558,18 +554,18 @@ void OpDispatchBuilder::AVX128_InsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstEle
if (Op->Src[1].IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Op->Src[1], GetGPROpSize(), Op->Flags);
auto Src2 = LoadSourceGPR_WithOpSize(Op, Op->Src[1], GetGPROpSize(), Op->Flags);
Result.Low = _VSToFGPRInsert(OpSize::i128Bit, DstElementSize, SrcSize, Src1.Low, Src2, false);
} else if (SrcSize != DstElementSize) {
// If the source is from memory but the Source size and destination size aren't the same,
// then it is more optimal to load in to a GPR and convert between GPR->FPR.
// ARM GPR->FPR conversion supports different size source and destinations while FPR->FPR doesn't.
auto Src2 = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags);
auto Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags);
Result.Low = _VSToFGPRInsert(DstSize, DstElementSize, SrcSize, Src1.Low, Src2, false);
} else {
// In the case of cvtsi2s{s,d} where the source and destination are the same size,
// then it is more optimal to load in to the FPR register directly and convert there.
auto Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
auto Src2 = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
// Always signed
Result.Low = _VSToFVectorInsert(DstSize, DstElementSize, DstElementSize, Src1.Low, Src2, false, false);
}
@@ -589,11 +585,11 @@ void OpDispatchBuilder::AVX128_CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSi
if (Op->Src[0].IsGPR()) {
Src = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, false);
} else {
Src.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcElementSize, Op->Flags);
Src.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[0], SrcElementSize, Op->Flags);
}
Ref Result = CVTFPR_To_GPRImpl(Op, Src.Low, SrcElementSize, HostRoundingMode);
StoreResult(GPRClass, Op, Result, OpSize::iInvalid);
StoreResultGPR(Op, Result);
}
void OpDispatchBuilder::AVX128_VANDN(OpcodeArgs) {
@@ -636,7 +632,7 @@ void OpDispatchBuilder::AVX128_UCOMISx(OpcodeArgs, IR::OpSize ElementSize) {
if (Op->Src[0].IsGPR()) {
Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, false);
} else {
Src2.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags);
Src2.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
}
Comiss(ElementSize, Src1.Low, Src2.Low);
@@ -653,7 +649,7 @@ void OpDispatchBuilder::AVX128_VectorScalarInsertALU(OpcodeArgs, FEXCore::IR::IR
if (Op->Src[1].IsGPR()) {
Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, false);
} else {
Src2.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], SrcSize, Op->Flags);
Src2.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[1], SrcSize, Op->Flags);
}
// If OpSize == ElementSize then it only does the lower scalar op
@@ -690,7 +686,7 @@ void OpDispatchBuilder::AVX128_InsertScalarFCMP(OpcodeArgs, IR::OpSize ElementSi
if (Op->Src[1].IsGPR()) {
Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, false);
} else {
Src2.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], SrcSize, Op->Flags);
Src2.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[1], SrcSize, Op->Flags);
}
const uint8_t CompType = Op->Src[2].Literal();
@@ -708,12 +704,12 @@ void OpDispatchBuilder::AVX128_MOVBetweenGPR_FPR(OpcodeArgs) {
RefPair Result {};
if (Op->Src[0].IsGPR()) {
// Loading from GPR and moving to Vector.
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], GetGPROpSize(), Op->Flags);
Ref Src = LoadSourceFPR_WithOpSize(Op, Op->Src[0], GetGPROpSize(), Op->Flags);
// zext to 128bit
Result.Low = _VCastFromGPR(OpSize::i128Bit, OpSizeFromSrc(Op), Src);
} else {
// Loading from Memory as a scalar. Zero extend
Result.Low = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Result.Low = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
}
Result.High = LoadZeroVector(OpSize::i128Bit);
@@ -726,11 +722,11 @@ void OpDispatchBuilder::AVX128_MOVBetweenGPR_FPR(OpcodeArgs) {
auto ElementSize = OpSizeFromDst(Op);
// Extract element from GPR. Zero extending in the process.
Src.Low = _VExtractToGPR(OpSizeFromSrc(Op), ElementSize, Src.Low, 0);
StoreResult(GPRClass, Op, Op->Dest, Src.Low, OpSize::iInvalid);
StoreResultGPR(Op, Op->Dest, Src.Low);
} else {
// Storing first element to memory.
Ref Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
_StoreMem(FPRClass, OpSizeFromDst(Op), Dest, Src.Low, OpSize::i8Bit);
Ref Dest = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.LoadData = false});
_StoreMemFPR(OpSizeFromDst(Op), Dest, Src.Low, OpSize::i8Bit);
}
}
}
@@ -758,7 +754,7 @@ void OpDispatchBuilder::AVX128_PExtr(OpcodeArgs, IR::OpSize ElementSize) {
const auto GPRSize = GetGPROpSize();
// Extract already zero extends the result.
Ref Result = _VExtractToGPR(OpSize::i128Bit, OverridenElementSize, Src.Low, Index);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Result, GPRSize, OpSize::iInvalid);
StoreResultGPR_WithOpSize(Op, Op->Dest, Result, GPRSize);
return;
}
@@ -779,7 +775,7 @@ void OpDispatchBuilder::AVX128_ExtendVectorElements(OpcodeArgs, IR::OpSize Eleme
const auto SrcSize = OpSizeFromSrc(Op);
const auto LoadSize = Is256Bit ? IR::SizeToOpSize(IR::OpSizeToSize(SrcSize) * 2) : SrcSize;
return LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], LoadSize, Op->Flags);
return LoadSourceFPR_WithOpSize(Op, Op->Src[0], LoadSize, Op->Flags);
}
};
@@ -868,7 +864,7 @@ void OpDispatchBuilder::AVX128_MOVMSK(OpcodeArgs, IR::OpSize ElementSize) {
auto GPRHigh = Mask8Byte(Src.High);
GPR = _Orlshl(OpSize::i64Bit, GPRLow, GPRHigh, 2);
}
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, GetGPROpSize(), OpSize::iInvalid);
StoreResultGPR_WithOpSize(Op, Op->Dest, GPR, GetGPROpSize());
}
void OpDispatchBuilder::AVX128_MOVMSKB(OpcodeArgs) {
@@ -897,7 +893,7 @@ void OpDispatchBuilder::AVX128_MOVMSKB(OpcodeArgs) {
Result = _Orlshl(OpSize::i64Bit, Result, ResultHigh, 16);
}
StoreResult(GPRClass, Op, Result, OpSize::iInvalid);
StoreResultGPR(Op, Result);
}
void OpDispatchBuilder::AVX128_PINSRImpl(OpcodeArgs, IR::OpSize ElementSize, const X86Tables::DecodedOperand& Src1Op,
@@ -910,7 +906,7 @@ void OpDispatchBuilder::AVX128_PINSRImpl(OpcodeArgs, IR::OpSize ElementSize, con
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
auto Src2 = LoadSourceGPR_WithOpSize(Op, Src2Op, GetGPROpSize(), Op->Flags);
Result.Low = _VInsGPR(OpSize::i128Bit, ElementSize, Index, Src1.Low, Src2);
} else {
// If loading from memory then we only load the element size
@@ -1047,7 +1043,7 @@ void OpDispatchBuilder::AVX128_InsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::O
// Then zero extends the top 128-bit.
const auto SrcSize = Op->Src[1].IsGPR() ? OpSize::i128Bit : SrcElementSize;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, false);
Ref Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Src2 = LoadSourceFPR_WithOpSize(Op, Op->Src[1], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Result = _VFToFScalarInsert(OpSize::i128Bit, DstElementSize, SrcElementSize, Src1.Low, Src2, false);
AVX128_StoreResult_WithOpSize(Op, Op->Dest, AVX128_Zext(Result));
@@ -1076,7 +1072,7 @@ void OpDispatchBuilder::AVX128_Vector_CVT_Float_To_Float(OpcodeArgs, IR::OpSize
} else {
// Handle 64-bit memory source.
// In the case of cvtps2pd xmm, m64.
Src.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], LoadSize, Op->Flags);
Src.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[0], LoadSize, Op->Flags);
}
RefPair Result {};
@@ -1154,7 +1150,7 @@ void OpDispatchBuilder::AVX128_Vector_CVT_Int_To_Float(OpcodeArgs, IR::OpSize Sr
// unnecessarily zero extend the vector. Otherwise, if
// memory, then we want to load the element size exactly.
const auto LoadSize = IR::SizeToOpSize(8 * (IR::OpSizeToSize(Size) / 16));
return RefPair {.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], LoadSize, Op->Flags)};
return RefPair {.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[0], LoadSize, Op->Flags)};
} else {
return AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, !Is128Bit);
}
@@ -1304,7 +1300,7 @@ void OpDispatchBuilder::AVX128_InsertScalarRound(OpcodeArgs, IR::OpSize ElementS
if (Op->Src[1].IsGPR()) {
Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, false);
} else {
Src2.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], SrcSize, Op->Flags);
Src2.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[1], SrcSize, Op->Flags);
}
// If OpSize == ElementSize then it only does the lower scalar op
@@ -1573,20 +1569,20 @@ void OpDispatchBuilder::AVX128_VMASKMOVImpl(OpcodeArgs, IR::OpSize ElementSize,
auto Address = MakeAddress(Op->Dest);
auto Data = AVX128_LoadSource_WithOpSize(Op, DataOp, Op->Flags, !Is128Bit);
_VStoreVectorMasked(OpSize::i128Bit, ElementSize, Mask.Low, Data.Low, Address, Invalid(), MEM_OFFSET_SXTX, 1);
_VStoreVectorMasked(OpSize::i128Bit, ElementSize, Mask.Low, Data.Low, Address, Invalid(), MemOffsetType::SXTX, 1);
if (!Is128Bit) {
_VStoreVectorMasked(OpSize::i128Bit, ElementSize, Mask.High, Data.High, Address, _InlineConstant(16), MEM_OFFSET_SXTX, 1);
_VStoreVectorMasked(OpSize::i128Bit, ElementSize, Mask.High, Data.High, Address, _InlineConstant(16), MemOffsetType::SXTX, 1);
}
} else {
auto Address = MakeAddress(DataOp);
RefPair Result {};
Result.Low = _VLoadVectorMasked(OpSize::i128Bit, ElementSize, Mask.Low, Address, Invalid(), MEM_OFFSET_SXTX, 1);
Result.Low = _VLoadVectorMasked(OpSize::i128Bit, ElementSize, Mask.Low, Address, Invalid(), MemOffsetType::SXTX, 1);
if (Is128Bit) {
Result.High = LoadZeroVector(OpSize::i128Bit);
} else {
Result.High = _VLoadVectorMasked(OpSize::i128Bit, ElementSize, Mask.High, Address, _InlineConstant(16), MEM_OFFSET_SXTX, 1);
Result.High = _VLoadVectorMasked(OpSize::i128Bit, ElementSize, Mask.High, Address, _InlineConstant(16), MemOffsetType::SXTX, 1);
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
}
@@ -1616,11 +1612,11 @@ void OpDispatchBuilder::AVX128_MASKMOV(OpcodeArgs) {
// RDI source (DS prefix by default)
auto MemDest = MakeSegmentAddress(X86State::REG_RDI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref XMMReg = _LoadMem(FPRClass, Size, MemDest, OpSize::i8Bit);
Ref XMMReg = _LoadMemFPR(Size, MemDest, OpSize::i8Bit);
// If the Mask element high bit is set then overwrite the element with the source, else keep the memory variant
XMMReg = _VBSL(Size, MaskSrc.Low, VectorSrc.Low, XMMReg);
_StoreMem(FPRClass, Size, MemDest, XMMReg, OpSize::i8Bit);
_StoreMemFPR(Size, MemDest, XMMReg, OpSize::i8Bit);
}
void OpDispatchBuilder::AVX128_VectorVariableBlend(OpcodeArgs, IR::OpSize ElementSize) {
@@ -1660,7 +1656,7 @@ void OpDispatchBuilder::AVX128_SaveAVXState(Ref MemBase) {
for (uint32_t i = 0; i < NumRegs; i += 2) {
RefPair Pair = LoadContextPair(OpSize::i128Bit, AVXHigh0Index + i);
_StoreMemPair(FPRClass, OpSize::i128Bit, Pair.Low, Pair.High, MemBase, i * 16 + 576);
_StoreMemPairFPR(OpSize::i128Bit, Pair.Low, Pair.High, MemBase, i * 16 + 576);
}
}
@@ -1668,7 +1664,7 @@ void OpDispatchBuilder::AVX128_RestoreAVXState(Ref MemBase) {
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
auto YMMHRegs = LoadMemPair(FPRClass, OpSize::i128Bit, MemBase, i * 16 + 576);
auto YMMHRegs = LoadMemPairFPR(OpSize::i128Bit, MemBase, i * 16 + 576);
AVX128_StoreXMMRegister(i, YMMHRegs.Low, true);
AVX128_StoreXMMRegister(i + 1, YMMHRegs.High, true);
@@ -1960,7 +1956,7 @@ void OpDispatchBuilder::AVX128_VFMAImpl(OpcodeArgs, IROps IROp, uint8_t Src1Idx,
}
void OpDispatchBuilder::AVX128_VFMAScalarImpl(OpcodeArgs, IROps IROp, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx) {
const auto SrcSize = OpSizeFromSrc(Op);
const OpSize ElementSize = Op->Flags & X86Tables::DecodeFlags::FLAG_OPTION_AVX_W ? OpSize::i64Bit : OpSize::i32Bit;
auto Dest = AVX128_LoadSource_WithOpSize(Op, Op->Dest, Op->Flags, false).Low;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, false).Low;
@@ -1968,13 +1964,13 @@ void OpDispatchBuilder::AVX128_VFMAScalarImpl(OpcodeArgs, IROps IROp, uint8_t Sr
if (Op->Src[1].IsGPR()) {
Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, false).Low;
} else {
Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], SrcSize, Op->Flags);
Src2 = LoadSourceFPR_WithOpSize(Op, Op->Src[1], ElementSize, Op->Flags);
}
Ref Sources[3] = {Dest, Src1, Src2};
DeriveOp(Result_Low, IROp,
_VFMLAScalarInsert(OpSize::i128Bit, SrcSize, Dest, Sources[Src1Idx - 1], Sources[Src2Idx - 1], Sources[AddendIdx - 1]));
_VFMLAScalarInsert(OpSize::i128Bit, ElementSize, Dest, Sources[Src1Idx - 1], Sources[Src2Idx - 1], Sources[AddendIdx - 1]));
AVX128_StoreResult_WithOpSize(Op, Op->Dest, AVX128_Zext(Result_Low));
}
@@ -2016,8 +2012,8 @@ void OpDispatchBuilder::AVX128_VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Sr
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
}
OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherImpl(OpSize Size, OpSize ElementLoadSize, OpSize AddrElementSize, RefPair Dest,
RefPair Mask, RefVSIB VSIB) {
OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherImpl(OpcodeArgs, OpSize Size, OpSize ElementLoadSize, OpSize AddrElementSize,
RefPair Dest, RefPair Mask, RefVSIB VSIB) {
LOGMAN_THROW_A_FMT(AddrElementSize == OpSize::i32Bit || AddrElementSize == OpSize::i64Bit, "Unknown address element size");
const auto Is128Bit = Size == OpSize::i128Bit;
@@ -2061,10 +2057,13 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherImpl(OpSize Size, O
}
}
const auto GPRSize = GetGPROpSize();
auto AddrSize = (Op->Flags & X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) != 0 ? (GPRSize >> 1) : GPRSize;
RefPair Result {};
///< Calculate the low-half.
Result.Low = _VLoadVectorGatherMasked(OpSize::i128Bit, ElementLoadSize, Dest.Low, Mask.Low, BaseAddr, VSIB.Low, VSIB.High,
AddrElementSize, VSIB.Scale, 0, 0);
AddrElementSize, VSIB.Scale, 0, 0, AddrSize);
if (Is128Bit) {
Result.High = LoadZeroVector(OpSize::i128Bit);
@@ -2101,7 +2100,7 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherImpl(OpSize Size, O
///< Calculate the high-half.
auto ResultHigh = _VLoadVectorGatherMasked(OpSize::i128Bit, ElementLoadSize, DestReg, MaskReg, BaseAddr, AddrAddressing.Low,
AddrAddressing.High, AddrElementSize, VSIB.Scale, DataElementOffset, IndexElementOffset);
AddrAddressing.High, AddrElementSize, VSIB.Scale, DataElementOffset, IndexElementOffset, AddrSize);
if (AddrElementSize == OpSize::i64Bit && ElementLoadSize == OpSize::i32Bit) {
// If we only fetched 128-bits worth of data then the upper-result is all zero.
@@ -2114,7 +2113,7 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherImpl(OpSize Size, O
return Result;
}
OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherQPSImpl(Ref Dest, Ref Mask, RefVSIB VSIB) {
OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherQPSImpl(OpcodeArgs, Ref Dest, Ref Mask, RefVSIB VSIB) {
///< BaseAddr doesn't need to exist, calculate that here.
Ref BaseAddr = VSIB.BaseAddr;
@@ -2142,8 +2141,11 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherQPSImpl(Ref Dest, R
RefPair Result {};
const auto GPRSize = GetGPROpSize();
auto AddrSize = (Op->Flags & X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) != 0 ? (GPRSize >> 1) : GPRSize;
///< Calculate the low-half.
Result.Low = _VLoadVectorGatherMaskedQPS(OpSize::i128Bit, OpSize::i32Bit, Dest, Mask, BaseAddr, VSIB.Low, VSIB.High, VSIB.Scale);
Result.Low = _VLoadVectorGatherMaskedQPS(OpSize::i128Bit, OpSize::i32Bit, Dest, Mask, BaseAddr, VSIB.Low, VSIB.High, VSIB.Scale, AddrSize);
Result.High = LoadZeroVector(OpSize::i128Bit);
if (VSIB.High == Invalid()) {
// Special case for only loading two floats.
@@ -2202,15 +2204,15 @@ void OpDispatchBuilder::AVX128_VPGATHER(OpcodeArgs, OpSize AddrElementSize) {
}
///< AddressElementSize is now OpSize::i64Bit
Result = AVX128_VPGatherQPSImpl(Dest.Low, Mask.Low, VSIBLow);
Result = AVX128_VPGatherQPSImpl(Op, Dest.Low, Mask.Low, VSIBLow);
if (NeedsHighAddrBytes) {
auto Res = AVX128_VPGatherQPSImpl(Dest.High, Mask.High, VSIBHigh);
auto Res = AVX128_VPGatherQPSImpl(Op, Dest.High, Mask.High, VSIBHigh);
Result.High = Res.Low;
}
} else if (AddrElementSize == OpSize::i64Bit && ElementLoadSize == OpSize::i32Bit) {
Result = AVX128_VPGatherQPSImpl(Dest.Low, Mask.Low, VSIB);
Result = AVX128_VPGatherQPSImpl(Op, Dest.Low, Mask.Low, VSIB);
} else {
Result = AVX128_VPGatherImpl(Size, ElementLoadSize, AddrElementSize, Dest, Mask, VSIB);
Result = AVX128_VPGatherImpl(Op, Size, ElementLoadSize, AddrElementSize, Dest, Mask, VSIB);
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
@@ -2234,7 +2236,7 @@ void OpDispatchBuilder::AVX128_VCVTPH2PS(OpcodeArgs) {
// In the event that a memory operand is used as the source operand,
// the access width will always be half the size of the destination vector width
// (i.e. 128-bit vector -> 64-bit mem, 256-bit vector -> 128-bit mem)
Src.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags);
Src.Low = LoadSourceFPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
}
RefPair Result {};
@@ -2289,7 +2291,7 @@ void OpDispatchBuilder::AVX128_VCVTPS2PH(OpcodeArgs) {
}
if (!Op->Dest.IsGPR()) {
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result.Low, StoreSize, OpSize::iInvalid);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result.Low, StoreSize);
} else {
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
}
@@ -23,8 +23,8 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
@@ -36,7 +36,7 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
@@ -44,15 +44,15 @@ void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref NewVec = _VExtr(OpSize::i128Bit, OpSize::i64Bit, Dest, Src, 1);
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
Ref Result = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, NewVec);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
@@ -60,8 +60,8 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// ARM SHA1 mostly matches x86 semantics, except the input and outputs are both flipped from elements 0,1,2,3 to 3,2,1,0.
auto Src1 = SHADataShuffle(Dest);
@@ -70,7 +70,7 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
// The result is swizzled differently than expected
auto Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
@@ -79,8 +79,8 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
return;
}
const uint64_t Imm8 = Op->Src[1].Literal() & 0b11;
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result {};
Ref ConstantVector {};
@@ -112,7 +112,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
@@ -120,12 +120,12 @@ void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
auto Result = _VSha256U0(Dest, Src);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
@@ -133,8 +133,8 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
auto Src1 = _VExtr(OpSize::i128Bit, OpSize::i32Bit, Dest, Dest, 3);
auto DupDst = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
@@ -142,7 +142,7 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
auto Result = _VSha256U1(Src1, Src2);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
@@ -150,8 +150,8 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
@@ -177,7 +177,7 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
auto B = _VSha256H2(EFGH, ABCD, Key);
auto Result = shuffle_abcd(A, B);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
@@ -185,9 +185,9 @@ void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESImc(Src);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
@@ -195,10 +195,10 @@ void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEnc(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
@@ -208,11 +208,11 @@ void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEnc(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
@@ -220,10 +220,10 @@ void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEncLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
@@ -233,11 +233,11 @@ void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEncLast(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
@@ -245,10 +245,10 @@ void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDec(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
@@ -258,11 +258,11 @@ void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDec(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
@@ -270,10 +270,10 @@ void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDecLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
@@ -283,15 +283,15 @@ void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDecLast(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const uint64_t RCON = Op->Src[1].Literal();
auto KeyGenSwizzle = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE);
@@ -305,7 +305,7 @@ void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
}
Ref Result = AESKeyGenAssistImpl(Op);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
@@ -313,12 +313,12 @@ void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Literal());
auto Res = _PCLMUL(OpSize::i128Bit, Dest, Src, Selector & 0b1'0001);
StoreResult(FPRClass, Op, Res, OpSize::iInvalid);
StoreResultFPR(Op, Res);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
@@ -328,12 +328,12 @@ void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
}
const auto DstSize = OpSizeFromDst(Op);
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Src1 = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Literal());
Ref Res = _PCLMUL(DstSize, Src1, Src2, Selector & 0b1'0001);
StoreResult(FPRClass, Op, Res, OpSize::iInvalid);
StoreResultFPR(Op, Res);
}
} // namespace FEXCore::IR
@@ -263,7 +263,7 @@ void OpDispatchBuilder::CalculateDeferredFlags() {
Ref OpDispatchBuilder::IncrementByCarry(OpSize OpSize, Ref Src) {
// If CF not inverted, we use .cc since the increment happens when the
// condition is false. If CF inverted, invert to use .cs. A bit mindbendy.
return _NZCVSelectIncrement(OpSize, {CFInverted ? COND_UGE : COND_ULT}, Src, Src);
return _NZCVSelectIncrement(OpSize, CFInverted ? CondClass::UGE : CondClass::ULT, Src, Src);
}
Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2) {
@@ -290,7 +290,7 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
// TODO: We can fold that second Bfe in (cmp uxth).
auto SelectCFInv = Select01(OpSize, CondClassType {COND_UGE}, Res, Src2PlusCF);
auto SelectCFInv = Select01(OpSize, CondClass::UGE, Res, Src2PlusCF);
SetNZ_ZeroCV(SrcSize, Res);
SetCFInverted(SelectCFInv);
@@ -324,7 +324,7 @@ Ref OpDispatchBuilder::CalculateFlags_SBB(IR::OpSize SrcSize, Ref Src1, Ref Src2
Res = Sub(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
auto SelectCFInv = Select01(OpSize, CondClassType {COND_UGE}, Src1, Src2PlusCF);
auto SelectCFInv = Select01(OpSize, CondClass::UGE, Src1, Src2PlusCF);
SetNZ_ZeroCV(SrcSize, Res);
SetCFInverted(SelectCFInv);
@@ -406,7 +406,7 @@ void OpDispatchBuilder::CalculateFlags_MUL(IR::OpSize SrcSize, Ref Res, Ref High
// If High = SignBit, then sets to nZCv. Else sets to nzcV. Since SF/ZF
// undefined, this does what we need after inverting carry.
auto Zero = _InlineConstant(0);
_CondSubNZCV(OpSize::i64Bit, Zero, Zero, CondClassType {COND_EQ}, 0x1 /* nzcV */);
_CondSubNZCV(OpSize::i64Bit, Zero, Zero, CondClass::EQ, 0x1 /* nzcV */);
CFInverted = true;
}
@@ -423,7 +423,7 @@ void OpDispatchBuilder::CalculateFlags_UMUL(Ref High) {
// If High = 0, then sets to nZCv. Else sets to nzcV. Since SF/ZF undefined,
// this does what we need.
_CondSubNZCV(Size, Zero, Zero, CondClassType {COND_EQ}, 0x1 /* nzcV */);
_CondSubNZCV(Size, Zero, Zero, CondClass::EQ, 0x1 /* nzcV */);
CFInverted = true;
}
@@ -151,6 +151,9 @@ constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_F2, 3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, false, 3>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_F2, 4), 4, &OpDispatchBuilder::NOPOp},
// GROUP 17
{OPD(FEXCore::X86Tables::TYPE_GROUP_17, PF_66, 0), 1, &OpDispatchBuilder::Extrq_imm},
// GROUP P
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_NONE, 0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, false, 1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_NONE, 1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, true, false, 1>},
@@ -145,7 +145,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
#ifndef _WIN32
// FEX reserved instructions
{0x37, 1, &OpDispatchBuilder::CallbackReturnOp},
{0x3E, 1, &OpDispatchBuilder::CallbackReturnOp},
{0x3F, 1, &OpDispatchBuilder::ThunkOp},
#endif
};
@@ -198,6 +198,8 @@ constexpr DispatchTableEntry OpDispatch_SecondaryRepNEModTables[] = {
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, true>},
{0x78, 1, &OpDispatchBuilder::Insertq_imm},
{0x79, 1, &OpDispatchBuilder::Insertq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i32Bit>},
@@ -256,6 +258,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x75, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i16Bit>},
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i32Bit>},
{0x78, 1, nullptr}, // GROUP 17
{0x79, 1, &OpDispatchBuilder::Extrq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i64Bit>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
File diff suppressed because it is too large. Load diff
@@ -17,7 +17,6 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/FPState.h>
#include <cmath>
#include <stddef.h>
#include <stdint.h>
@@ -28,7 +27,7 @@ class OrderedNode;
Ref OpDispatchBuilder::GetX87Top() {
// Yes, we are storing 3 bits in a single flag register.
// Deal with it
return _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
return _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
void OpDispatchBuilder::SetX87FTW(Ref FTW) {
@@ -52,24 +51,22 @@ void OpDispatchBuilder::SetX87FTW(Ref FTW) {
FTW = _Orlshr(OpSize::i32Bit, FTW, FTW, 4);
// ...and that's it. StoreContext implicitly does the final masking.
_StoreContext(OpSize::i8Bit, GPRClass, FTW, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
_StoreContextGPR(OpSize::i8Bit, FTW, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
}
void OpDispatchBuilder::SetX87Top(Ref Value) {
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
// Float LoaD operation with memory operand
void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], Width, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
Ref ConvertedData = Data;
// Convert to 80bit float
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
ConvertedData = _F80CVTTo(Data, ReadWidth);
ConvertedData = _F80CVTTo(Data, Width);
}
_PushStack(ConvertedData, Data, ReadWidth, true);
_PushStack(ConvertedData, Data, Width);
}
// Float LoaD operation with memory operand
@@ -79,27 +76,27 @@ void OpDispatchBuilder::FLDFromStack(OpcodeArgs) {
void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Data, OpSize::i128Bit, true);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
Ref converted = _F80BCDStore(_ReadStackValue(0));
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
StoreResultFPR_WithOpSize(Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
_PopStackDestroy();
}
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::i128Bit, true);
_PushStack(Data, Data, OpSize::f80Bit);
}
void OpDispatchBuilder::FILD(OpcodeArgs) {
const auto ReadWidth = OpSizeFromSrc(Op);
// Read from memory
Ref Data = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
Ref Data = LoadSourceGPR_WithOpSize(Op, Op->Src[0], ReadWidth, Op->Flags);
// Sign extend to 64bits
if (ReadWidth != OpSize::i64Bit) {
@@ -112,27 +109,28 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
// Extract sign and make integer absolute
auto zero = Constant(0);
_SubNZCV(OpSize::i64Bit, Data, zero);
auto sign = _NZCVSelect(OpSize::i64Bit, CondClassType {COND_SLT}, Constant(0x8000), zero);
auto absolute = _Neg(OpSize::i64Bit, Data, CondClassType {COND_MI});
auto sign = _NZCVSelect(OpSize::i64Bit, CondClass::SLT, Constant(0x8000), zero);
auto absolute = _Neg(OpSize::i64Bit, Data, CondClass::MI);
// left justify the absolute integer
auto shift = Sub(OpSize::i64Bit, Constant(63), _FindMSB(IR::OpSize::i64Bit, absolute));
auto shifted = _Lshl(OpSize::i64Bit, absolute, shift);
auto adjusted_exponent = Sub(OpSize::i64Bit, Constant(0x3fff + 63), shift);
auto zeroed_exponent = _Select(COND_EQ, absolute, zero, zero, adjusted_exponent);
auto zeroed_exponent = _Select(OpSize::i64Bit, OpSize::i64Bit, CondClass::EQ, absolute, zero, zero, adjusted_exponent);
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VLoadTwoGPRs(shifted, upper);
_PushStack(ConvertedData, Data, ReadWidth, false);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Width == OpSize::i32Bit || Width == OpSize::i64Bit || Width == OpSize::f80Bit, "Invalid store width for FST");
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::f80Bit;
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale, /*Float=*/true);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
_PopStackDestroy();
@@ -166,12 +164,12 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
Ref IsSpecial = _NZCVSelect01({COND_EQ});
Ref IsSpecial = _NZCVSelect01(CondClass::EQ);
// For overflow detection, check if exponent indicates a value >= 2^15
// Biased exponent for 2^15 is 0x3fff + 15 = 0x400e
SubWithFlags(OpSize::i64Bit, Exponent, 0x400e);
Ref IsOverflow = _NZCVSelect01({COND_UGE});
Ref IsOverflow = _NZCVSelect01(CondClass::UGE);
// Set Invalid Operation flag if overflow or special value
Ref InvalidFlag = _Or(OpSize::i64Bit, IsSpecial, IsOverflow);
@@ -180,7 +178,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Data = _F80CVTInt(Size, Data, Truncate);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -206,10 +204,10 @@ void OpDispatchBuilder::FADD(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispa
// We have one memory argument
Ref Arg {};
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTToInt(Arg, Width);
} else {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTTo(Arg, Width);
}
@@ -236,10 +234,10 @@ void OpDispatchBuilder::FMUL(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispa
// We have one memory argument
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTTo(arg, Width);
}
@@ -273,10 +271,10 @@ void OpDispatchBuilder::FDIV(OpcodeArgs, IR::OpSize Width, bool Integer, bool Re
// We have one memory argument
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTTo(arg, Width);
}
@@ -314,10 +312,10 @@ void OpDispatchBuilder::FSUB(OpcodeArgs, IR::OpSize Width, bool Integer, bool Re
// We have one memory argument
Ref Arg {};
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTToInt(Arg, Width);
} else {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTTo(Arg, Width);
}
@@ -340,7 +338,7 @@ Ref OpDispatchBuilder::GetX87FTW_Helper() {
// bytes, we use the well-known bit twiddling algorithm:
//
// https://graphics.stanford.edu/~seander/bithacks.html#InterleaveBMN
Ref X = _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
Ref X = _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
X = _Orlshl(OpSize::i32Bit, X, X, 4);
X = _And(OpSize::i32Bit, X, Constant(0x0f0f0f0f));
X = _Orlshl(OpSize::i32Bit, X, X, 2);
@@ -381,41 +379,41 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
_SyncStackToSlow();
const auto Size = OpSizeFromSrc(Op);
Ref Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
Ref Mem = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
{
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
auto FCW = _LoadContextGPR(OpSize::i16Bit, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMemGPR(Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMemGPR(Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MemOffsetType::SXTX, 1); }
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MemOffsetType::SXTX, 1);
}
}
@@ -441,20 +439,20 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
_StackForceSlow();
const auto Size = OpSizeFromSrc(Op);
Ref Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, {.LoadData = false});
Ref Mem = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFCW = _LoadMemGPR(OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
Ref MemLocation = Add(OpSize::i64Bit, Mem, IR::OpSizeToSize(Size) * 1);
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
auto NewFSW = _LoadMemGPR(Size, MemLocation, Size);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
Ref MemLocation = Add(OpSize::i64Bit, Mem, IR::OpSizeToSize(Size) * 2);
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
SetX87FTW(_LoadMemGPR(Size, MemLocation, Size));
}
}
@@ -483,61 +481,61 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
Ref Top = GetX87Top();
{
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
auto FCW = _LoadContextGPR(OpSize::i16Bit, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMemGPR(Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMemGPR(Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MemOffsetType::SXTX, 1); }
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MemOffsetType::SXTX, 1);
}
auto SevenConst = Constant(7);
const auto LoadSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
Ref data = _LoadContextFPRIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
_StoreMem(FPRClass, OpSize::i128Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMemFPR(OpSize::i128Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
Top = _And(OpSize::i32Bit, Add(OpSize::i32Bit, Top, 1), SevenConst);
}
// The final st(7) needs a bit of special handling here
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
Ref data = _LoadContextFPRIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
_StoreMem(FPRClass, OpSize::i64Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMemFPR(OpSize::i64Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
auto topBytes = _VDupElement(OpSize::i128Bit, OpSize::i16Bit, data, 4);
_StoreMem(FPRClass, OpSize::i16Bit, topBytes, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMemFPR(OpSize::i16Bit, topBytes, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MemOffsetType::SXTX, 1);
// reset to default
FNINIT(Op);
@@ -548,8 +546,8 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFCW = _LoadMemGPR(OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
if (ReducedPrecisionMode) {
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
@@ -561,11 +559,11 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
_SetRoundingMode(roundingMode, false, roundingMode);
}
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MemOffsetType::SXTX, 1);
Ref Top = ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1));
}
auto SevenConst = Constant(7);
@@ -574,14 +572,14 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
Ref Mask = _VLoadTwoGPRs(low, high);
const auto StoreSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMem(FPRClass, OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMemFPR(OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
// Mask off the top bits
Reg = _VAnd(OpSize::i128Bit, OpSize::i128Bit, Reg, Mask);
if (ReducedPrecisionMode) {
// Convert to double precision
Reg = _F80CVT(OpSize::i64Bit, Reg);
}
_StoreContextIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
_StoreContextFPRIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
Top = _And(OpSize::i32Bit, Add(OpSize::i32Bit, Top, 1), SevenConst);
}
@@ -590,19 +588,19 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
Ref Reg = _LoadMem(FPRClass, OpSize::i64Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref RegHigh = _LoadMem(FPRClass, OpSize::i16Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMemFPR(OpSize::i64Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
Ref RegHigh = _LoadMemFPR(OpSize::i16Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MemOffsetType::SXTX, 1);
Reg = _VInsElement(OpSize::i128Bit, OpSize::i16Bit, 4, 0, Reg, RegHigh);
if (ReducedPrecisionMode) {
Reg = _F80CVT(OpSize::i64Bit, Reg); // Convert to double precision
}
_StoreContextIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
_StoreContextFPRIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
}
// Load / Store Control Word
void OpDispatchBuilder::X87FSTCW(OpcodeArgs) {
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
StoreResult(GPRClass, Op, FCW, OpSize::iInvalid);
auto FCW = _LoadContextGPR(OpSize::i16Bit, offsetof(FEXCore::Core::CPUState, FCW));
StoreResultGPR(Op, FCW);
}
void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
@@ -610,8 +608,8 @@ void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
// to switch for now to slow mode whenever these are manually changed.
// Remove the next line and try DF_04.asm in fast path.
_StackForceSlow();
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
Ref NewFCW = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
void OpDispatchBuilder::FXCH(OpcodeArgs) {
@@ -647,10 +645,10 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
// Memory arg
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
b = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
b = _F80CVTTo(arg, Width);
}
} else {
@@ -766,7 +764,7 @@ Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
void OpDispatchBuilder::X87FNSTSW(OpcodeArgs) {
Ref TopValue = _SyncStackToSlow();
Ref StatusWord = ReconstructFSW_Helper(TopValue);
StoreResult(GPRClass, Op, StatusWord, OpSize::iInvalid);
StoreResultGPR(Op, StatusWord);
}
void OpDispatchBuilder::FNCLEX(OpcodeArgs) {
@@ -785,12 +783,12 @@ void OpDispatchBuilder::FNINIT(OpcodeArgs) {
// Init FCW to 0x037F
auto NewFCW = Constant(0x037F);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
// Set top to zero
SetX87Top(Zero);
// Tags all get marked as invalid
_StoreContext(OpSize::i8Bit, GPRClass, Zero, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
_StoreContextGPR(OpSize::i8Bit, Zero, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
// Reinits the simulated stack
_InitStack();
@@ -863,7 +861,7 @@ void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
auto TopValid = _StackValidTag(0);
// In the case of top being invalid then C3:C2:C0 is 0b101
auto C3 = Select01(OpSize::i32Bit, CondClassType {COND_NEQ}, TopValid, Constant(1));
auto C3 = Select01(OpSize::i32Bit, CondClass::NEQ, TopValid, Constant(1));
auto C2 = TopValid;
auto C0 = C3; // Mirror C3 until something other than zero is supported
@@ -878,8 +876,8 @@ void OpDispatchBuilder::X87FXTRACT(OpcodeArgs) {
_PopStackDestroy();
auto Exp = _F80XTRACT_EXP(Top);
auto Sig = _F80XTRACT_SIG(Top);
_PushStack(Exp, Exp, OpSize::f80Bit, true);
_PushStack(Sig, Sig, OpSize::f80Bit, true);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -29,38 +29,37 @@ void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
auto NewFCW = _LoadMemGPR(OpSize::i16Bit, Mem, OpSize::i16Bit);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size)), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size)), Size, MemOffsetType::SXTX, 1);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1));
}
}
void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
_StackForceSlow();
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Ref NewFCW = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
// F64 ops
// Float load op with memory operand
void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], Width, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
// Convert to 64bit float
Ref ConvertedData = Data;
if (Width == OpSize::i32Bit) {
@@ -68,39 +67,39 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
} else if (Width == OpSize::f80Bit) {
ConvertedData = _F80CVT(OpSize::i64Bit, Data);
}
_PushStack(ConvertedData, Data, ReadWidth, true);
_PushStack(ConvertedData, Data, Width);
}
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
// Read from memory
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Data, OpSize::i64Bit, true);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
Ref converted = _F80CVTTo(_ReadStackValue(0), OpSize::i64Bit);
converted = _F80BCDStore(converted);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
StoreResultFPR_WithOpSize(Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
_PopStackDestroy();
}
void OpDispatchBuilder::FLDF64_Const(OpcodeArgs, uint64_t Num) {
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(Num));
_PushStack(Data, Data, OpSize::i64Bit, true);
_PushStack(Data, Data, OpSize::i64Bit);
}
void OpDispatchBuilder::FILDF64(OpcodeArgs) {
const auto ReadWidth = OpSizeFromSrc(Op);
// Read from memory
Ref Data = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
Ref Data = LoadSourceGPR_WithOpSize(Op, Op->Src[0], ReadWidth, Op->Flags);
if (ReadWidth == OpSize::i16Bit) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
auto ConvertedData = _Float_FromGPR_S(OpSize::i64Bit, ReadWidth == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, Data);
_PushStack(ConvertedData, Data, ReadWidth, false);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
@@ -112,7 +111,7 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -138,16 +137,16 @@ void OpDispatchBuilder::FADDF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpDi
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
} else {
FEX_UNREACHABLE;
}
@@ -176,16 +175,16 @@ void OpDispatchBuilder::FMULF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpDi
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
} else {
FEX_UNREACHABLE;
}
@@ -228,16 +227,16 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs, IR::OpSize Width, bool Integer, bool
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
Arg = _Sbfe(OpSize::i64Bit, 16, 0, Arg);
}
Arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, Arg);
} else if (Width == OpSize::i32Bit) {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, Arg);
} else if (Width == OpSize::i64Bit) {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
@@ -285,16 +284,16 @@ void OpDispatchBuilder::FSUBF64(OpcodeArgs, IR::OpSize Width, bool Integer, bool
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
@@ -332,16 +331,16 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpD
} else if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
// Memory arg
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
b = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
b = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
b = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
@@ -393,11 +392,11 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
SaveNZCV();
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, ExpZV, ExpNZV);
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
_PopStackDestroy();
_PushStack(Exp, Exp, OpSize::i64Bit, true);
_PushStack(Sig, Sig, OpSize::i64Bit, true);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -1,87 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
desc: Guest-side assembly helpers used by the backends
$end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#include <cstring>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
X86GeneratedCode::X86GeneratedCode() {
#ifdef _WIN32
// No need to allocate anything in this config.
#else
// Allocate a page for our emulated guest
CodePtr = AllocateGuestCodeSpace(CODE_SIZE);
constexpr std::array<uint8_t, 2> SignalReturnCode = {
0x0F, 0x37, // CALLBACKRET FEX Instruction
};
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
memcpy(reinterpret_cast<void*>(CallbackReturn), SignalReturnCode.data(), SignalReturnCode.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
#endif
}
X86GeneratedCode::~X86GeneratedCode() {
#ifndef _WIN32
FEXCore::Allocator::VirtualFree(CodePtr, CODE_SIZE);
#endif
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
#ifndef _WIN32
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
return FEXCore::Allocator::VirtualAlloc(Size);
}
// First 64bit page
constexpr uintptr_t LOCATION_MAX = 0x1'0000'0000;
// 32bit mode
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void* Ptr = ::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED && reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
::munmap(Ptr, Size);
continue;
}
if (Ptr != MAP_FAILED) {
return Ptr;
}
}
// Can't do anything about this
// Here's hoping the application doesn't use signals
return MAP_FAILED;
#else
return nullptr;
#endif
}
} // namespace FEXCore
@@ -1,25 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
$end_info$
*/
#pragma once
#include <stddef.h>
#include <stdint.h>
namespace FEXCore {
class X86GeneratedCode final {
public:
X86GeneratedCode();
~X86GeneratedCode();
uint64_t CallbackReturn {};
private:
void* CodePtr {};
void* AllocateGuestCodeSpace(size_t Size);
};
} // namespace FEXCore
@@ -100,10 +100,11 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0x34, 1, X86InstInfo{"SYSENTER", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x35, 1, X86InstInfo{"SYSEXIT", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x36, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x37, 1, X86InstInfo{"GETSEC", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x38, 1, X86InstInfo{"", TYPE_0F38_TABLE, FLAGS_NO_OVERLAY, 0}},
{0x39, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x3A, 1, X86InstInfo{"", TYPE_0F3A_TABLE, FLAGS_NO_OVERLAY, 0}},
{0x3B, 4, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x3B, 3, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x40, 1, X86InstInfo{"CMOVO", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0x41, 1, X86InstInfo{"CMOVNO", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
@@ -299,7 +300,7 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
// FEX reserved instructions
// Unused x86 encoding instruction.
{0x37, 1, X86InstInfo{"CALLBACKRET", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0}},
{0x3E, 1, X86InstInfo{"CALLBACKRET", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0}},
// This was originally used by VIA to jump to its alternative instruction set. Used for OP_THUNK
{0x3F, 1, X86InstInfo{"ALTINST", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0}},
@@ -442,7 +443,7 @@ constexpr std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps = []() c
{0x70, 1, X86InstInfo{"PSHUFLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{0x71, 3, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0}},
{0x74, 4, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{0x78, 1, X86InstInfo{"INSERTQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS,2}},
{0x78, 1, X86InstInfo{"INSERTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS,2}},
{0x79, 1, X86InstInfo{"INSERTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 0}},
{0x7A, 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{0x7C, 1, X86InstInfo{"HADDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
@@ -7,6 +7,7 @@ $end_info$
#pragma once
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
@@ -559,21 +560,6 @@ constexpr static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86T
}
};
template<typename OpcodeType>
static inline void LateInitCopyTable(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *OtherLocal, size_t OtherTableSize) {
for (size_t j = 0; j < OtherTableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &OtherOp = OtherLocal[j];
auto OtherOpNum = OtherOp.first;
X86InstInfo const &OtherInfo = OtherOp.Info;
for (uint32_t i = 0; i < OtherOp.second; ++i) {
X86InstInfo &FinalOp = FinalTable[OtherOpNum + i];
if (FinalOp.Type == TYPE_COPY_OTHER) {
FinalOp = OtherInfo;
}
}
}
}
template<typename OpcodeType>
constexpr static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
for (size_t j = 0; j < TableSize; ++j) {
@@ -604,12 +590,6 @@ constexpr static inline void GenerateX87Table(X86InstInfo *FinalTable, X86Tables
}
};
}
FEX_DEFINE_ENUM_FMT_PASSTHROUGH(FEXCore::X86Tables::DecodedOperand::OpType);
template <>
struct fmt::formatter<FEXCore::X86Tables::DecodedOperand::OpType> : formatter<uint32_t> {
template <typename FormatContext>
auto format(FEXCore::X86Tables::DecodedOperand::OpType type, FormatContext& ctx) const {
return fmt::formatter<uint32_t>::format(static_cast<uint32_t>(type), ctx);
}
};
} // namespace FEXCore::X86Tables
+10 -149
View File
@@ -20,7 +20,6 @@
namespace FEXCore::IR {
class OrderedNode;
class RegisterAllocationPass;
/**
* @brief The IROp_Header is an dynamically sized array
@@ -61,24 +60,7 @@ struct NodeID final {
Value = 0;
}
[[nodiscard]] friend constexpr bool operator==(NodeID, NodeID) noexcept = default;
[[nodiscard]]
friend constexpr bool operator<(NodeID lhs, NodeID rhs) noexcept {
return lhs.Value < rhs.Value;
}
[[nodiscard]]
friend constexpr bool operator>(NodeID lhs, NodeID rhs) noexcept {
return operator<(rhs, lhs);
}
[[nodiscard]]
friend constexpr bool operator<=(NodeID lhs, NodeID rhs) noexcept {
return !operator>(lhs, rhs);
}
[[nodiscard]]
friend constexpr bool operator>=(NodeID lhs, NodeID rhs) noexcept {
return !operator<(lhs, rhs);
}
[[nodiscard]] constexpr auto operator<=>(const NodeID&) const noexcept = default;
friend std::ostream& operator<<(std::ostream& out, NodeID ID) {
out << ID.Value;
@@ -431,94 +413,6 @@ static_assert(sizeof(OrderedNode) == (sizeof(OrderedNodeHeader) + 2 * sizeof(uin
// };
using Ref = OrderedNode*;
struct FEX_PACKED RegisterClassType final {
using value_type = uint32_t;
value_type Val;
[[nodiscard]] constexpr operator value_type() const {
return Val;
}
[[nodiscard]]
friend constexpr bool operator==(const RegisterClassType&, const RegisterClassType&) = default;
};
struct FEX_PACKED CondClassType final {
uint8_t Val;
[[nodiscard]] constexpr operator uint8_t() const {
return Val;
}
[[nodiscard]]
friend constexpr bool operator==(const CondClassType&, const CondClassType&) = default;
};
struct FEX_PACKED MemOffsetType final {
uint8_t Val;
[[nodiscard]] constexpr operator uint8_t() const {
return Val;
}
[[nodiscard]]
friend constexpr bool operator==(const MemOffsetType&, const MemOffsetType&) = default;
};
struct FEX_PACKED TypeDefinition final {
uint16_t Val;
[[nodiscard]] constexpr operator uint16_t() const {
return Val;
}
[[nodiscard]]
static constexpr TypeDefinition Create(uint8_t Bytes) {
TypeDefinition Type {};
Type.Val = Bytes << 8;
return Type;
}
[[nodiscard]]
static constexpr TypeDefinition Create(uint8_t Bytes, uint8_t Elements) {
TypeDefinition Type {};
Type.Val = (Bytes << 8) | (Elements & 255);
return Type;
}
[[nodiscard]]
constexpr uint8_t Bytes() const {
return Val >> 8;
}
[[nodiscard]]
constexpr uint8_t Elements() const {
return Val & 255;
}
[[nodiscard]]
friend constexpr bool operator==(const TypeDefinition&, const TypeDefinition&) = default;
};
static_assert(std::is_trivially_copyable_v<TypeDefinition>);
struct FEX_PACKED FenceType final {
using value_type = uint8_t;
value_type Val;
[[nodiscard]] constexpr operator value_type() const {
return Val;
}
[[nodiscard]]
friend constexpr bool operator==(const FenceType&, const FenceType&) = default;
};
struct FEX_PACKED RoundType final {
uint8_t Val;
[[nodiscard]] constexpr operator uint8_t() const {
return Val;
}
[[nodiscard]]
friend constexpr bool operator==(const RoundType&, const RoundType&) = default;
};
class NodeIterator;
/* This iterator can be used to step though nodes.
* Due to how our IR is laid out, this can be used to either step
* though the CodeBlocks or though the code within a single block.
@@ -782,6 +676,15 @@ inline NodeID NodeWrapperBase<Type>::ID() const {
bool IsBlockExit(FEXCore::IR::IROps Op);
void Dump(fextl::stringstream* out, const IRListView* IR);
constexpr auto format_as(FEXCore::IR::NodeID ID) {
return ID.Value;
}
FEX_DEFINE_ENUM_FMT_PASSTHROUGH(FEXCore::IR::FenceType)
FEX_DEFINE_ENUM_FMT_PASSTHROUGH(FEXCore::IR::MemOffsetType)
FEX_DEFINE_ENUM_FMT_PASSTHROUGH(FEXCore::IR::OpSize)
FEX_DEFINE_ENUM_FMT_PASSTHROUGH(FEXCore::IR::RegClass)
} // namespace FEXCore::IR
template<>
@@ -790,45 +693,3 @@ struct std::hash<FEXCore::IR::NodeID> {
return std::hash<FEXCore::IR::NodeID::value_type> {}(ID.Value);
}
};
template<>
struct fmt::formatter<FEXCore::IR::NodeID> : fmt::formatter<FEXCore::IR::NodeID::value_type> {
using Base = fmt::formatter<FEXCore::IR::NodeID::value_type>;
// Pass-through the underlying value, so IDs can
// be formatted like any integral value.
template<typename FormatContext>
auto format(const FEXCore::IR::NodeID& ID, FormatContext& ctx) const {
return Base::format(ID.Value, ctx);
}
};
template<>
struct fmt::formatter<FEXCore::IR::RegisterClassType> : fmt::formatter<FEXCore::IR::RegisterClassType::value_type> {
using Base = fmt::formatter<FEXCore::IR::RegisterClassType::value_type>;
template<typename FormatContext>
auto format(const FEXCore::IR::RegisterClassType& Class, FormatContext& ctx) const {
return Base::format(Class.Val, ctx);
}
};
template<>
struct fmt::formatter<FEXCore::IR::FenceType> : fmt::formatter<FEXCore::IR::FenceType::value_type> {
using Base = fmt::formatter<FEXCore::IR::FenceType::value_type>;
template<typename FormatContext>
auto format(const FEXCore::IR::FenceType& Fence, FormatContext& ctx) const {
return Base::format(Fence.Val, ctx);
}
};
template<>
struct fmt::formatter<FEXCore::IR::OpSize> : fmt::formatter<std::underlying_type_t<FEXCore::IR::OpSize>> {
using Base = fmt::formatter<std::underlying_type_t<FEXCore::IR::OpSize>>;
template<typename FormatContext>
auto format(const FEXCore::IR::OpSize& OpSize, FormatContext& ctx) const {
return Base::format(FEXCore::ToUnderlying(OpSize), ctx);
}
};
+117 -142
View File
@@ -52,80 +52,68 @@
" * These are validations that can't be automatically inferred and need to be hand-written",
""
],
"Enums": {
"class CondClass : uint8_t": [
"EQ = 0,",
"NEQ = 1,",
"UGE = 2,",
"ULT = 3,",
"MI = 4,",
"PL = 5,",
"VS = 6,",
"VC = 7,",
"UGT = 8,",
"ULE = 9,",
"SGE = 10,",
"SLT = 11,",
"SGT = 12,",
"SLE = 13,",
"TSTZ = 14, /* bit test zero */",
"TSTNZ = 15, /* bit test nonzero */",
"",
"FLU = 16, /* float less or unordered */",
"FGE = 17, /* float greater or equal */",
"FLEU = 18, /* float less or equal or unordered */",
"FGT = 19, /* float greater */",
"FU = 20, /* float unordered */",
"FNU = 21, /* float not unordered */",
"",
"AL = 32, /* always */"
],
"class FenceType : uint8_t": [
"Load = 0,",
"Store = 1,",
"LoadStore = 2,",
"Inst = 3,"
],
"class MemOffsetType : uint8_t": [
"SXTX = 0,",
"UXTW = 1,",
"SXTW = 2,"
],
"class RegClass : uint32_t": [
"Invalid = 0,",
"GPR = 1,",
"GPRFixed = 2,",
"FPR = 3,",
"FPRFixed = 4,",
"Complex = 5,"
],
"class RoundMode : uint8_t": [
"Nearest = 0,",
"NegInfinity = 1,",
"PosInfinity = 2,",
"TowardsZero = 3, /* Truncate */",
"Host = 4,"
]
},
"Defines": [
"constexpr uint8_t COND_EQ = 0",
"constexpr uint8_t COND_NEQ = 1",
"constexpr uint8_t COND_UGE = 2",
"constexpr uint8_t COND_ULT = 3",
"constexpr uint8_t COND_MI = 4",
"constexpr uint8_t COND_PL = 5",
"constexpr uint8_t COND_VS = 6",
"constexpr uint8_t COND_VC = 7",
"constexpr uint8_t COND_UGT = 8",
"constexpr uint8_t COND_ULE = 9",
"constexpr uint8_t COND_SGE = 10",
"constexpr uint8_t COND_SLT = 11",
"constexpr uint8_t COND_SGT = 12",
"constexpr uint8_t COND_SLE = 13",
"constexpr uint8_t COND_TSTZ = 14 /* bit test zero */",
"constexpr uint8_t COND_TSTNZ = 15 /* bit test nonzero */",
"constexpr uint8_t COND_FLU = 16 /* float less or unordred */",
"constexpr uint8_t COND_FGE = 17 /* float greater or equal */",
"constexpr uint8_t COND_FLEU = 18 /* float less or equal or unordred */",
"constexpr uint8_t COND_FGT = 19 /* float greater */",
"constexpr uint8_t COND_FU = 20 /* float unordred */",
"constexpr uint8_t COND_FNU = 21 /* float not unordred */",
"constexpr uint8_t COND_AL = 32 /* always */",
"constexpr FEXCore::IR::RegisterClassType InvalidClass {0}",
"constexpr FEXCore::IR::RegisterClassType GPRClass {1}",
"constexpr FEXCore::IR::RegisterClassType GPRFixedClass {2}",
"constexpr FEXCore::IR::RegisterClassType FPRClass {3}",
"constexpr FEXCore::IR::RegisterClassType FPRFixedClass {4}",
"constexpr FEXCore::IR::RegisterClassType ComplexClass {5}",
"constexpr uint8_t NumClasses {6}",
"",
"constexpr FEXCore::IR::TypeDefinition i8 {TypeDefinition::Create(1, 0)}",
"constexpr FEXCore::IR::TypeDefinition i16 {TypeDefinition::Create(2, 0)}",
"constexpr FEXCore::IR::TypeDefinition i32 {TypeDefinition::Create(4, 0)}",
"constexpr FEXCore::IR::TypeDefinition i64 {TypeDefinition::Create(8, 0)}",
"constexpr FEXCore::IR::TypeDefinition i128 {TypeDefinition::Create(16, 0)}",
"",
"constexpr FEXCore::IR::TypeDefinition i8v8 {TypeDefinition::Create(1, 8)}",
"constexpr FEXCore::IR::TypeDefinition i8v16 {TypeDefinition::Create(1, 16)}",
"constexpr FEXCore::IR::TypeDefinition i16v4 {TypeDefinition::Create(2, 4)}",
"constexpr FEXCore::IR::TypeDefinition i16v8 {TypeDefinition::Create(2, 8)}",
"constexpr FEXCore::IR::TypeDefinition i32v2 {TypeDefinition::Create(4, 2)}",
"constexpr FEXCore::IR::TypeDefinition i32v4 {TypeDefinition::Create(4, 4)}",
"constexpr FEXCore::IR::TypeDefinition i64v2 {TypeDefinition::Create(8, 2)}",
"",
"constexpr uint8_t FCMP_FLAG_EQ = 0",
"constexpr uint8_t FCMP_FLAG_LT = 1",
"constexpr uint8_t FCMP_FLAG_UNORDERED = 2",
"constexpr FEXCore::IR::FenceType Fence_Load {0}",
"constexpr FEXCore::IR::FenceType Fence_Store {1}",
"constexpr FEXCore::IR::FenceType Fence_LoadStore {2}",
"constexpr FEXCore::IR::FenceType Fence_Inst {3}",
"constexpr uint8_t ROUND_MODE_NEAREST = 0",
"constexpr uint8_t ROUND_MODE_NEGATIVE_INFINITY = 1",
"constexpr uint8_t ROUND_MODE_POSITIVE_INFINITY = 2",
"constexpr uint8_t ROUND_MODE_TOWARDS_ZERO = 3",
"constexpr uint8_t ROUND_MODE_FLUSH_TO_ZERO = 1 << 2",
"constexpr FEXCore::IR::RoundType Round_Nearest {ROUND_MODE_NEAREST}",
"constexpr FEXCore::IR::RoundType Round_Negative_Infinity {ROUND_MODE_NEGATIVE_INFINITY}",
"constexpr FEXCore::IR::RoundType Round_Positive_Infinity {ROUND_MODE_POSITIVE_INFINITY}",
"constexpr FEXCore::IR::RoundType Round_Towards_Zero {ROUND_MODE_TOWARDS_ZERO} /* Truncate */",
"constexpr FEXCore::IR::RoundType Round_Host {ROUND_MODE_TOWARDS_ZERO + 1}",
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_SXTX {0}",
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_UXTW {1}",
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_SXTW {2}",
"struct BreakDefinition {",
" uint16_t ErrorRegister;",
" uint8_t Signal;",
@@ -148,13 +136,12 @@
"GPR": "OrderedNode*",
"FPR": "OrderedNode*",
"FenceType": "FenceType",
"RegisterClass": "RegisterClassType",
"CondClass": "CondClassType",
"SyscallFlags": "FEXCore::IR::SyscallFlags",
"RegisterClass": "RegClass",
"CondClass": "CondClass",
"SHA256Sum": "SHA256Sum",
"MemOffsetType": "MemOffsetType",
"BreakDefinition": "BreakDefinition",
"RoundType": "RoundType",
"RoundType": "RoundMode",
"FloatCompareOp": "FloatCompareOp",
"NamedVectorConstant": "FEXCore::IR::NamedVectorConstant",
"IndexNamedVectorConstant": "FEXCore::IR::IndexNamedVectorConstant",
@@ -307,7 +294,7 @@
"HasSideEffects": true,
"RAOverride": "0"
},
"CondJump SSA:$Cmp1, SSA:$Cmp2, SSA:$TrueBlock, SSA:$FalseBlock, CondClass:$Cond{{COND_NEQ}}, OpSize:$CompareSize{OpSize::iInvalid}, i1:$FromNZCV{false}": {
"CondJump SSA:$Cmp1, SSA:$Cmp2, SSA:$TrueBlock, SSA:$FalseBlock, CondClass:$Cond{CondClass::NEQ}, OpSize:$CompareSize{OpSize::iInvalid}, i1:$FromNZCV{false}": {
"Inline": ["", "AddSub"],
"HasSideEffects": true,
"RAOverride": "2"
@@ -326,25 +313,13 @@
"CallbackReturn": {
"HasSideEffects": true
},
"GPR = Syscall GPR:$SyscallID, GPR:$Arg0, GPR:$Arg1, GPR:$Arg2, GPR:$Arg3, GPR:$Arg4, GPR:$Arg5, SyscallFlags:$Flags": {
"GPR = Syscall GPR:$SyscallID, GPR:$Arg0, GPR:$Arg1, GPR:$Arg2, GPR:$Arg3, GPR:$Arg4, GPR:$Arg5": {
"HasSideEffects": true,
"Desc": ["Dispatches a guest syscall through to the SyscallHandler class"
],
"DestSize": "OpSize::i64Bit"
},
"GPR = InlineSyscall GPR:$Arg0, GPR:$Arg1, GPR:$Arg2, GPR:$Arg3, GPR:$Arg4, GPR:$Arg5, i32:$HostSyscallNumber, SyscallFlags:$Flags": {
"HasSideEffects": true,
"Desc": ["Dispatches a guest syscall directly to the host syscall interface,",
"bypassing the SyscallHandler class used by Syscall.",
"This has significantly less overhead than Syscall, which needs to save JIT state first.",
"Can only be used for syscalls that match across architecture,",
"such as gettid (matches on x86/x86-64/Arm64)."
],
"DestSize": "OpSize::i64Bit"
},
"Thunk GPR:$ArgPtr, SHA256Sum:$ThunkNameHash": {
"HasSideEffects": true
},
@@ -364,20 +339,6 @@
"GPR = Copy GPR:$Source": {
"Desc": ["GPR copy, generated by RA to split live ranges"],
"DestSize": "OpSize::i64Bit"
},
"GPR = Swap1 GPR:$A, GPR:$B": {
"Desc": ["GPR swap part 1, generated by RA. Returns value of first source.",
"Destination must be second GPR."],
"DestSize": "OpSize::i64Bit"
},
"GPR = Swap2": {
"Desc": ["GPR swap part 2, generated by RA. Returns source source.",
"Must immediately succeed Swap1 with no intervening instructions",
"Kludge to workaround single destination restriction on IR",
"Hopefully temporary"],
"DestSize": "OpSize::i64Bit"
}
},
"StaticRA": {
@@ -423,8 +384,8 @@
],
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == GPRClass",
"($Class == RegClass::GPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == RegClass::FPR",
"($Class == RegClass::FPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == RegClass::GPR",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't LoadContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't LoadContext to XMM\""
]
@@ -437,8 +398,8 @@
"HasSideEffects": true,
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == GPRClass",
"($Class == RegClass::GPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == RegClass::FPR",
"($Class == RegClass::FPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == RegClass::GPR",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't LoadContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't LoadContext to XMM\""
]
@@ -454,8 +415,8 @@
"HasSideEffects": true,
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == GPRClass",
"($Class == RegClass::GPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == RegClass::FPR",
"($Class == RegClass::FPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == RegClass::GPR",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't StoreContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't StoreContext to XMM\""
]
@@ -472,8 +433,8 @@
"EmitValidation": [
"WalkFindRegClass($Value1) == $Class",
"WalkFindRegClass($Value2) == $Class",
"($Class == GPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == GPRClass",
"($Class == RegClass::GPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == RegClass::FPR",
"($Class == RegClass::FPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == RegClass::GPR",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't StoreContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't StoreContext to XMM\""
]
@@ -485,8 +446,8 @@
],
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == GPRClass",
"($Class == RegClass::GPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == RegClass::FPR",
"($Class == RegClass::FPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == RegClass::GPR",
"!($BaseOffset >= offsetof(Core::CPUState, gregs[0]) && $BaseOffset < offsetof(Core::CPUState, gregs[16])) && \"Can't LoadContextIndexed to GPR\"",
"!($BaseOffset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $BaseOffset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't LoadContextIndexed to XMM\""
]
@@ -499,8 +460,8 @@
"DestSize": "ByteSize",
"EmitValidation": [
"WalkFindRegClass($Value) == $Class",
"($Class == GPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == GPRClass",
"($Class == RegClass::GPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit)) || $Class == RegClass::FPR",
"($Class == RegClass::FPR && (#ByteSize == IR::OpSize::i8Bit || #ByteSize == IR::OpSize::i16Bit || #ByteSize == IR::OpSize::i32Bit || #ByteSize == IR::OpSize::i64Bit || #ByteSize == IR::OpSize::i128Bit || #ByteSize == IR::OpSize::i256Bit)) || $Class == RegClass::GPR",
"!($BaseOffset >= offsetof(Core::CPUState, gregs[0]) && $BaseOffset < offsetof(Core::CPUState, gregs[16])) && \"Can't StoreContextIndexed to GPR\"",
"!($BaseOffset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $BaseOffset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't StoreContextIndexed to XMM\""
]
@@ -630,7 +591,7 @@
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = VLoadVectorGatherMasked OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Incoming, FPR:$Mask, GPR:$AddrBase, FPR:$VectorIndexLow, FPR:$VectorIndexHigh, OpSize:$VectorIndexElementSize, u8:$OffsetScale, u8:$DataElementOffsetStart, u8:$IndexElementOffsetStart": {
"FPR = VLoadVectorGatherMasked OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Incoming, FPR:$Mask, GPR:$AddrBase, FPR:$VectorIndexLow, FPR:$VectorIndexHigh, OpSize:$VectorIndexElementSize, u8:$OffsetScale, u8:$DataElementOffsetStart, u8:$IndexElementOffsetStart, OpSize:$AddrSize": {
"Desc": [
"Does a masked load similar to VPGATHERD* where the upper bit of each element",
"determines whether or not that element will be loaded from memory.",
@@ -644,7 +605,7 @@
"$VectorIndexElementSize == OpSize::i32Bit || $VectorIndexElementSize == OpSize::i64Bit"
]
},
"FPR = VLoadVectorGatherMaskedQPS OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Incoming, FPR:$MaskReg, GPR:$AddrBase, FPR:$VectorIndexLow, FPR:$VectorIndexHigh, u8:$OffsetScale": {
"FPR = VLoadVectorGatherMaskedQPS OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Incoming, FPR:$MaskReg, GPR:$AddrBase, FPR:$VectorIndexLow, FPR:$VectorIndexHigh, u8:$OffsetScale, OpSize:$AddrSize": {
"Desc": [
"Does a masked load similar to VPGATHERQPS where the upper bit of each element",
"determines whether or not that element will be loaded from memory.",
@@ -759,9 +720,10 @@
},
"Fence FenceType:$Fence": {
"Desc": ["Does a memory fence operation of the desired type",
"Fence_Load: Ensures load memory operations are serialized",
"Fence_Store: Ensures store memory operations are serialized",
"Fence_LoadStore: Ensures loads and store memory operations are serialized",
"FenceType::Load: Ensures load memory operations are serialized",
"FenceType::Store: Ensures store memory operations are serialized",
"FenceType::LoadStore: Ensures loads and store memory operations are serialized",
"FenceType::Inst: Instruction barrier. Ensures all instructions after this point will be explicitly fetched",
"Ensures the memory operations are globally visible"
],
"HasSideEffects": true
@@ -842,16 +804,6 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AtomicXor OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer xor",
"IR layout must match Fetch-variant, otherwise DCE IR optimization breaks!"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i8Bit || Size == FEXCore::IR::OpSize::i16Bit || Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AtomicSwap OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer swap"
@@ -1006,7 +958,7 @@
"DestSize": "OpSize::i64Bit"
},
"GPR = Neg OpSize:#Size, GPR:$Src, CondClass:$Cond{{COND_AL}}": {
"GPR = Neg OpSize:#Size, GPR:$Src, CondClass:$Cond{CondClass::AL}": {
"Desc": ["Integer negation, with optional predication",
"Dest = Cond ? -Src : Src",
"Will truncate to 64 or 32bits"
@@ -1084,6 +1036,13 @@
"Size == FEXCore::IR::OpSize::i16Bit || Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = Rbit OpSize:#Size, GPR:$Src": {
"Desc": ["Reverses the bit order of the register"],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = Add OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": [ "Integer Add",
"Will truncate to 64 or 32bits"
@@ -1551,6 +1510,15 @@
"ResultSize == FEXCore::IR::OpSize::i32Bit || ResultSize == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = MaskGenerateFromBitWidth GPR:$BitWidth": {
"Desc": ["Generates a bit mask from with a value from [0, 63]",
"0 is special cased to full-mask",
"Special operation for SSE4a bitmask generation."
],
"DestSize": "FEXCore::IR::OpSize::i64Bit",
"ImplicitFlagClobber": true
},
"GPR = Extr OpSize:#Size, GPR:$Upper, GPR:$Lower, u8:$LSB": {
"Desc": ["Concats the two GPRs to create a value that is the size of the full two GPRs",
"It then extracts a bitfield width that size of a GPR from the LSB",
@@ -2159,22 +2127,34 @@
"FPR = VAnd OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
"ElementSize": "ElementSize",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VAndn OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
"ElementSize": "ElementSize",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VOr OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
"ElementSize": "ElementSize",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VXor OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
"ElementSize": "ElementSize",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VUQAdd OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
@@ -2828,17 +2808,13 @@
"X87": true,
"HasSideEffects": true
},
"PushStack FPR:$X80Src, SSA:$OriginalValue, OpSize:$LoadSize, i1:$Float": {
"PushStack FPR:$X80Src, FPR:$OriginalValue, OpSize:$LoadSize": {
"Desc": [
"Pushes the provided X80Src source on to the x87 stack.",
"Tracks OriginalValue as the original value of X80Src.",
"Tracks OriginalValue as the original value of X80Src. OriginalValue can be Invalid() in which case no tracking is done.",
"Opsize is 128bit for F80 values, 64-bit for low precision.",
"LoadSize the original load size, i.e. of size of OriginalValue.",
"Float: 80-bit, 64-bit, 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"EmitValidation": [
"WalkFindRegClass($OriginalValue) == FPRClass || WalkFindRegClass($OriginalValue) == GPRClass"
"Float: 80-bit, 64-bit, 32-bit"
],
"HasSideEffects": true,
"X87": true
@@ -2850,13 +2826,12 @@
"HasSideEffects": true,
"X87": true
},
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale, i1:$Float": {
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale": {
"Desc": [
"Takes the top value off the x87 stack and stores it to memory.",
"SourceSize is 128bit for F80 values, 64-bit for low precision.",
"StoreSize is the store size for conversion:",
"Float: 80-bit, 64-bit, or 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
"Float: 80-bit, 64-bit, or 32-bit"
],
"HasSideEffects": true,
"X87": true
+136 -107
View File
@@ -38,8 +38,8 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, uint64_t Arg)
*out << fextl::fmt::format("#{:#x}", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, CondClassType Arg) {
if (Arg == COND_AL) {
static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg) {
if (Arg == CondClass::AL) {
*out << "ALWAYS";
return;
}
@@ -48,7 +48,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, CondClassType
"UGT", "ULE", "SGE", "SLT", "SGT", "SLE", "TSTZ", "TSTNZ",
"FLU", "FGE", "FLEU", "FGT", "FU", "FNU"};
*out << CondNames[Arg];
*out << CondNames[FEXCore::ToUnderlying(Arg)];
}
static void PrintArg(fextl::stringstream* out, const IRListView*, MemOffsetType Arg) {
@@ -58,39 +58,39 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, MemOffsetType
"SXTW",
};
*out << Names[Arg];
*out << Names[FEXCore::ToUnderlying(Arg)];
}
static void PrintArg(fextl::stringstream* out, const IRListView*, RegisterClassType Arg) {
if (Arg == GPRClass.Val) {
*out << "GPR";
} else if (Arg == GPRFixedClass.Val) {
*out << "GPRFixed";
} else if (Arg == FPRClass.Val) {
*out << "FPR";
} else if (Arg == FPRFixedClass.Val) {
*out << "FPRFixed";
} else {
*out << "Unknown Registerclass " << Arg;
}
static void PrintArg(fextl::stringstream* out, const IRListView*, RegClass Arg) {
*out << [Arg] {
switch (Arg) {
case RegClass::Invalid: return "Invalid";
case RegClass::GPR: return "GPR";
case RegClass::GPRFixed: return "GPRFixed";
case RegClass::FPR: return "FPR";
case RegClass::FPRFixed: return "FPRFixed";
case RegClass::Complex: return "Complex";
}
return "<Unknown RegClass Type>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNodeWrapper Arg) {
if (Arg.IsImmediate()) {
auto PhyReg = PhysicalRegister(Arg);
switch (PhyReg.Class) {
case FEXCore::IR::GPRClass.Val: *out << "r"; break;
case FEXCore::IR::GPRFixedClass.Val: *out << "R"; break;
case FEXCore::IR::FPRClass.Val: *out << "v"; break;
case FEXCore::IR::FPRFixedClass.Val: *out << "V"; break;
case FEXCore::IR::ComplexClass.Val: *out << "c"; break;
case FEXCore::IR::InvalidClass.Val: *out << "invalid"; break;
switch (PhyReg.AsRegClass()) {
case RegClass::GPR: *out << "r"; break;
case RegClass::GPRFixed: *out << "R"; break;
case RegClass::FPR: *out << "v"; break;
case RegClass::FPRFixed: *out << "V"; break;
case RegClass::Complex: *out << "c"; break;
case RegClass::Invalid: *out << "invalid"; break;
default: *out << "unknown"; break;
}
if (PhyReg.Class != FEXCore::IR::InvalidClass.Val) {
*out << std::dec << (uint32_t)PhyReg.Reg;
if (PhyReg.AsRegClass() != RegClass::Invalid) {
*out << std::dec << uint32_t(PhyReg.Reg);
}
return;
@@ -124,41 +124,32 @@ static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNode
}
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::FenceType Arg) {
if (Arg == IR::Fence_Load) {
*out << "Loads";
} else if (Arg == IR::Fence_Store) {
*out << "Stores";
} else if (Arg == IR::Fence_LoadStore) {
*out << "LoadStores";
} else {
*out << "<Unknown Fence Type>";
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FenceType Arg) {
*out << [Arg] {
switch (Arg) {
case FenceType::Load: return "Loads";
case FenceType::Store: return "Stores";
case FenceType::LoadStore: return "LoadStores";
case FenceType::Inst: return "Instruction";
}
return "<Unknown Fence Type>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::RoundType Arg) {
switch (Arg) {
case FEXCore::IR::Round_Nearest: *out << "Nearest"; break;
case FEXCore::IR::Round_Negative_Infinity: *out << "-Inf"; break;
case FEXCore::IR::Round_Positive_Infinity: *out << "+Inf"; break;
case FEXCore::IR::Round_Towards_Zero: *out << "Towards Zero"; break;
case FEXCore::IR::Round_Host: *out << "Host"; break;
default: *out << "<Unknown Round Type>"; break;
}
static void PrintArg(fextl::stringstream* out, const IRListView*, RoundMode Arg) {
*out << [Arg] {
switch (Arg) {
case RoundMode::Nearest: return "Nearest";
case RoundMode::NegInfinity: return "-Inf";
case RoundMode::PosInfinity: return "+Inf";
case RoundMode::TowardsZero: return "Towards Zero";
case RoundMode::Host: return "Host";
}
return "<Unknown Round Type>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::SyscallFlags Arg) {
switch (Arg) {
case FEXCore::IR::SyscallFlags::DEFAULT: *out << "Default"; break;
case FEXCore::IR::SyscallFlags::OPTIMIZETHROUGH: *out << "Optimize Through"; break;
case FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY: *out << "No Sync State on Entry"; break;
case FEXCore::IR::SyscallFlags::NORETURN: *out << "No Return"; break;
case FEXCore::IR::SyscallFlags::NOSIDEEFFECTS: *out << "No Side Effects"; break;
default: *out << "<Unknown Round Type>"; break;
}
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::NamedVectorConstant Arg) {
static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorConstant Arg) {
*out << [Arg] {
// clang-format off
switch (Arg) {
@@ -186,6 +177,22 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::N
return "movmskps_shift";
case NamedVectorConstant::NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE:
return "aeskeygenassist_swizzle";
case NamedVectorConstant::NAMED_VECTOR_BLENDPS_0110B:
return "blendps_0110b";
case NamedVectorConstant::NAMED_VECTOR_BLENDPS_0111B:
return "blendps_0111b";
case NamedVectorConstant::NAMED_VECTOR_BLENDPS_1001B:
return "blendps_1001b";
case NamedVectorConstant::NAMED_VECTOR_BLENDPS_1011B:
return "blendps_1011b";
case NamedVectorConstant::NAMED_VECTOR_BLENDPS_1101B:
return "blendps_1101b";
case NamedVectorConstant::NAMED_VECTOR_BLENDPS_1110B:
return "blendps_1110b";
case NamedVectorConstant::NAMED_VECTOR_MOVMASKB:
return "movmaskb";
case NamedVectorConstant::NAMED_VECTOR_MOVMASKB_UPPER:
return "movmaskb_upper";
case NamedVectorConstant::NAMED_VECTOR_ZERO:
return "vectorzero";
case NamedVectorConstant::NAMED_VECTOR_X87_ONE:
@@ -216,9 +223,20 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::N
return "cvtmax_i32";
case NamedVectorConstant::NAMED_VECTOR_CVTMAX_I64:
return "cvtmax_i64";
default:
return "<Unknown Named Vector Constant>";
case NamedVectorConstant::NAMED_VECTOR_F80_SIGN_MASK:
return "f80_sign_mask";
case NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K0:
return "sha1rnds_k0";
case NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K1:
return "sha1rnds_k1";
case NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K2:
return "sha1rnds_k2";
case NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K3:
return "sha1rnds_k3";
case NamedVectorConstant::NAMED_VECTOR_MAX:
return "<Programming Error: Printing MAX value>";
}
return "<Unknown Named Vector Constant>";
// clang-format on
}();
}
@@ -241,36 +259,43 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, IndexNamedVect
return "dppd_mask";
case IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PBLENDW:
return "pblendw";
default:
return "<Unknown Indexed Named Vector Constant>";
case INDEXED_NAMED_VECTOR_MAX:
return "<Programming Error: Printing MAX value>";
}
return "<Unknown Indexed Named Vector Constant>";
// clang-format on
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::OpSize Arg) {
switch (Arg) {
case OpSize::i8Bit: *out << "i8"; break;
case OpSize::i16Bit: *out << "i16"; break;
case OpSize::i32Bit: *out << "i32"; break;
case OpSize::i64Bit: *out << "i64"; break;
case OpSize::i128Bit: *out << "i128"; break;
case OpSize::i256Bit: *out << "i256"; break;
case OpSize::f80Bit: *out << "f80"; break;
default: *out << "<Unknown OpSize Type>"; break;
}
static void PrintArg(fextl::stringstream* out, const IRListView*, OpSize Arg) {
*out << [Arg] {
switch (Arg) {
case OpSize::iUnsized: return "Unsized";
case OpSize::i8Bit: return "i8";
case OpSize::i16Bit: return "i16";
case OpSize::i32Bit: return "i32";
case OpSize::i64Bit: return "i64";
case OpSize::f80Bit: return "f80";
case OpSize::i128Bit: return "i128";
case OpSize::i256Bit: return "i256";
case OpSize::iInvalid: return "Invalid";
}
return "<Unknown OpSize Type>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::FloatCompareOp Arg) {
switch (Arg) {
case FloatCompareOp::EQ: *out << "FEQ"; break;
case FloatCompareOp::LT: *out << "FLT"; break;
case FloatCompareOp::LE: *out << "FLE"; break;
case FloatCompareOp::UNO: *out << "UNO"; break;
case FloatCompareOp::NEQ: *out << "NEQ"; break;
case FloatCompareOp::ORD: *out << "ORD"; break;
default: *out << "<Unknown OpSize Type>"; break;
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FloatCompareOp Arg) {
*out << [Arg] {
switch (Arg) {
case FloatCompareOp::EQ: return "FEQ";
case FloatCompareOp::LT: return "FLT";
case FloatCompareOp::LE: return "FLE";
case FloatCompareOp::UNO: return "UNO";
case FloatCompareOp::NEQ: return "NEQ";
case FloatCompareOp::ORD: return "ORD";
}
return "<Unknown FloatCompareOp Type>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::BreakDefinition Arg) {
@@ -280,23 +305,28 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::B
*out << static_cast<uint32_t>(Arg.si_code) << "}";
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::ShiftType Arg) {
switch (Arg) {
case ShiftType::LSL: *out << "LSL"; break;
case ShiftType::LSR: *out << "LSR"; break;
case ShiftType::ASR: *out << "ASR"; break;
case ShiftType::ROR: *out << "ROR"; break;
default: *out << "<Unknown Shift Type>"; break;
}
static void PrintArg(fextl::stringstream* out, const IRListView*, ShiftType Arg) {
*out << [Arg] {
switch (Arg) {
case ShiftType::LSL: return "LSL";
case ShiftType::LSR: return "LSR";
case ShiftType::ASR: return "ASR";
case ShiftType::ROR: return "ROR";
}
return "<Unknown Shift Type>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::BranchHint Arg) {
switch (Arg) {
case BranchHint::None: *out << "None"; break;
case BranchHint::Call: *out << "Call"; break;
case BranchHint::Return: *out << "Return"; break;
default: *out << "<Unknown Branch Hint>"; break;
}
static void PrintArg(fextl::stringstream* out, const IRListView*, BranchHint Arg) {
*out << [Arg] {
switch (Arg) {
case BranchHint::None: return "None";
case BranchHint::Call: return "Call";
case BranchHint::Return: return "Return";
case BranchHint::CheckTF: return "CheckTF";
}
return "<Unknown Branch Hint>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, const std::array<uint8_t, 0x10>& Arg) {
@@ -323,8 +353,7 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
AddIndent();
*out << "(%" << IR->GetID(BlockNode) << ") "
<< "CodeBlock ";
*out << "(%" << IR->GetID(BlockNode) << ") " << "CodeBlock ";
*out << "%" << BlockIROp->Begin.ID() << ", ";
*out << "%" << BlockIROp->Last.ID() << std::endl;
@@ -353,17 +382,17 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
auto PhyReg = PhysicalRegister(CodeNode);
if (!PhyReg.IsInvalid()) {
switch (PhyReg.Class) {
case FEXCore::IR::GPRClass.Val: *out << "(r"; break;
case FEXCore::IR::GPRFixedClass.Val: *out << "(R"; break;
case FEXCore::IR::FPRClass.Val: *out << "(v"; break;
case FEXCore::IR::FPRFixedClass.Val: *out << "(V"; break;
case FEXCore::IR::ComplexClass.Val: *out << "(complex"; break;
case FEXCore::IR::InvalidClass.Val: *out << "(invalid"; break;
switch (PhyReg.AsRegClass()) {
case RegClass::GPR: *out << "(r"; break;
case RegClass::GPRFixed: *out << "(R"; break;
case RegClass::FPR: *out << "(v"; break;
case RegClass::FPRFixed: *out << "(V"; break;
case RegClass::Complex: *out << "(complex"; break;
case RegClass::Invalid: *out << "(invalid"; break;
default: *out << "(unknown"; break;
}
if (PhyReg.Class != FEXCore::IR::InvalidClass.Val) {
*out << std::dec << (uint32_t)PhyReg.Reg << ")";
if (PhyReg.AsRegClass() != RegClass::Invalid) {
*out << std::dec << uint32_t(PhyReg.Reg) << ")";
} else {
*out << ")";
}
+7 -7
View File
@@ -33,14 +33,14 @@ bool IsBlockExit(FEXCore::IR::IROps Op) {
}
}
FEXCore::IR::RegisterClassType IREmitter::WalkFindRegClass(Ref Node) {
RegClass IREmitter::WalkFindRegClass(Ref Node) {
auto Class = GetOpRegClass(Node);
switch (Class) {
case GPRClass:
case FPRClass:
case GPRFixedClass:
case FPRFixedClass:
case InvalidClass: return Class;
case RegClass::GPR:
case RegClass::FPR:
case RegClass::GPRFixed:
case RegClass::FPRFixed:
case RegClass::Invalid: return Class;
default: break;
}
@@ -82,7 +82,7 @@ FEXCore::IR::RegisterClassType IREmitter::WalkFindRegClass(Ref Node) {
}
default: LOGMAN_MSG_A_FMT("Unhandled op type: {} {} in argument class validation", ToUnderlying(IROp->Op), GetOpName(Node)); break;
}
return InvalidClass;
return RegClass::Invalid;
}
void IREmitter::ResetWorkingList() {
+68 -18
View File
@@ -46,12 +46,12 @@ public:
*
* @{ */
FEXCore::IR::RegisterClassType WalkFindRegClass(Ref Node);
RegClass WalkFindRegClass(Ref Node);
// These inlining helpers are used by IRDefines.inc so define first.
Ref InlineMem(OpSize Size, Ref Offset, MemOffsetType OffsetType, uint8_t& OffsetScale, bool TSO = false) {
uint64_t Imm {};
if (OffsetType != MEM_OFFSET_SXTX || !IsValueConstant(WrapNode(Offset), &Imm)) {
if (OffsetType != MemOffsetType::SXTX || !IsValueConstant(WrapNode(Offset), &Imm)) {
return Offset;
}
@@ -108,36 +108,86 @@ public:
IRPair<IROp_Jump> _Jump() {
return _Jump(InvalidNode);
}
IRPair<IROp_CondJump> _CondJump(Ref ssa0, CondClassType cond = {COND_NEQ}) {
IRPair<IROp_CondJump> _CondJump(Ref ssa0, CondClass cond = CondClass::NEQ) {
return _CondJump(ssa0, _Constant(0), InvalidNode, InvalidNode, cond, GetOpSize(ssa0));
}
IRPair<IROp_CondJump> _CondJump(Ref ssa0, Ref ssa1, Ref ssa2, CondClassType cond = {COND_NEQ}) {
IRPair<IROp_CondJump> _CondJump(Ref ssa0, Ref ssa1, Ref ssa2, CondClass cond = CondClass::NEQ) {
return _CondJump(ssa0, _Constant(0), ssa1, ssa2, cond, GetOpSize(ssa0));
}
// TODO: Work to remove this implicit sized Select implementation.
IRPair<IROp_Select> _Select(uint8_t Cond, Ref ssa0, Ref ssa1, Ref ssa2, Ref ssa3, IR::OpSize CompareSize = OpSize::iUnsized) {
if (CompareSize == OpSize::iUnsized) {
CompareSize = std::max(OpSize::i32Bit, std::max(GetOpSize(ssa0), GetOpSize(ssa1)));
}
return _Select(std::max(OpSize::i32Bit, std::max(GetOpSize(ssa2), GetOpSize(ssa3))), CompareSize, CondClassType {Cond}, ssa0, ssa1, ssa2, ssa3);
IRPair<IROp_LoadContext> _LoadContextGPR(OpSize ByteSize, uint32_t Offset) {
return _LoadContext(ByteSize, RegClass::GPR, Offset);
}
IRPair<IROp_LoadMem> _LoadMem(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref ssa0, IR::OpSize Align = OpSize::i8Bit) {
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
IRPair<IROp_LoadContext> _LoadContextFPR(OpSize ByteSize, uint32_t Offset) {
return _LoadContext(ByteSize, RegClass::FPR, Offset);
}
IRPair<IROp_StoreMem> _StoreMem(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, Ref Addr, Ref Value, IR::OpSize Align = OpSize::i8Bit) {
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
IRPair<IROp_StoreContext> _StoreContextGPR(OpSize ByteSize, Ref Value, uint32_t Offset) {
return _StoreContext(ByteSize, RegClass::GPR, Value, Offset);
}
IRPair<IROp_StoreContext> _StoreContextFPR(OpSize ByteSize, Ref Value, uint32_t Offset) {
return _StoreContext(ByteSize, RegClass::FPR, Value, Offset);
}
IRPair<IROp_Select> Select01(FEXCore::IR::OpSize CompareSize, CondClassType Cond, OrderedNode* Cmp1, OrderedNode* Cmp2) {
IRPair<IROp_LoadContextIndexed> _LoadContextGPRIndexed(Ref Index, OpSize ByteSize, uint32_t BaseOffset, uint32_t Stride) {
return _LoadContextIndexed(Index, ByteSize, BaseOffset, Stride, RegClass::GPR);
}
IRPair<IROp_LoadContextIndexed> _LoadContextFPRIndexed(Ref Index, OpSize ByteSize, uint32_t BaseOffset, uint32_t Stride) {
return _LoadContextIndexed(Index, ByteSize, BaseOffset, Stride, RegClass::FPR);
}
IRPair<IROp_StoreContextIndexed> _StoreContextGPRIndexed(Ref Value, Ref Index, OpSize ByteSize, uint32_t BaseOffset, uint32_t Stride) {
return _StoreContextIndexed(Value, Index, ByteSize, BaseOffset, Stride, RegClass::GPR);
}
IRPair<IROp_StoreContextIndexed> _StoreContextFPRIndexed(Ref Value, Ref Index, OpSize ByteSize, uint32_t BaseOffset, uint32_t Stride) {
return _StoreContextIndexed(Value, Index, ByteSize, BaseOffset, Stride, RegClass::FPR);
}
IRPair<IROp_LoadMem> _LoadMem(RegClass Class, OpSize Size, Ref ssa0, OpSize Align = OpSize::i8Bit) {
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MemOffsetType::SXTX, 1);
}
IRPair<IROp_LoadMem> _LoadMemGPR(OpSize Size, Ref ssa0, OpSize Align = OpSize::i8Bit) {
return _LoadMem(RegClass::GPR, Size, ssa0, Invalid(), Align, MemOffsetType::SXTX, 1);
}
IRPair<IROp_LoadMem> _LoadMemGPR(OpSize Size, Ref Addr, Ref Offset, OpSize Align, MemOffsetType OffsetType, uint8_t OffsetScale) {
return _LoadMem(RegClass::GPR, Size, Addr, Offset, Align, OffsetType, OffsetScale);
}
IRPair<IROp_LoadMem> _LoadMemFPR(OpSize Size, Ref ssa0, OpSize Align = OpSize::i8Bit) {
return _LoadMem(RegClass::FPR, Size, ssa0, Invalid(), Align, MemOffsetType::SXTX, 1);
}
IRPair<IROp_LoadMem> _LoadMemFPR(OpSize Size, Ref Addr, Ref Offset, OpSize Align, MemOffsetType OffsetType, uint8_t OffsetScale) {
return _LoadMem(RegClass::FPR, Size, Addr, Offset, Align, OffsetType, OffsetScale);
}
IRPair<IROp_StoreMem> _StoreMem(RegClass Class, OpSize Size, Ref Addr, Ref Value, OpSize Align = OpSize::i8Bit) {
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MemOffsetType::SXTX, 1);
}
IRPair<IROp_StoreMem> _StoreMemGPR(OpSize Size, Ref Addr, Ref Value, OpSize Align = OpSize::i8Bit) {
return _StoreMem(RegClass::GPR, Size, Value, Addr, Invalid(), Align, MemOffsetType::SXTX, 1);
}
IRPair<IROp_StoreMem> _StoreMemGPR(OpSize Size, Ref Value, Ref Addr, Ref Offset, OpSize Align, MemOffsetType OffsetType, uint8_t OffsetScale) {
return _StoreMem(RegClass::GPR, Size, Value, Addr, Offset, Align, OffsetType, OffsetScale);
}
IRPair<IROp_StoreMem> _StoreMemFPR(OpSize Size, Ref Addr, Ref Value, OpSize Align = OpSize::i8Bit) {
return _StoreMem(RegClass::FPR, Size, Value, Addr, Invalid(), Align, MemOffsetType::SXTX, 1);
}
IRPair<IROp_StoreMem> _StoreMemFPR(OpSize Size, Ref Value, Ref Addr, Ref Offset, OpSize Align, MemOffsetType OffsetType, uint8_t OffsetScale) {
return _StoreMem(RegClass::FPR, Size, Value, Addr, Offset, Align, OffsetType, OffsetScale);
}
IRPair<IROp_StoreMemPair> _StoreMemPairGPR(OpSize Size, Ref Value1, Ref Value2, Ref Addr, uint32_t Offset) {
return _StoreMemPair(RegClass::GPR, Size, Value1, Value2, Addr, Offset);
}
IRPair<IROp_StoreMemPair> _StoreMemPairFPR(OpSize Size, Ref Value1, Ref Value2, Ref Addr, uint32_t Offset) {
return _StoreMemPair(RegClass::FPR, Size, Value1, Value2, Addr, Offset);
}
IRPair<IROp_Select> Select01(FEXCore::IR::OpSize CompareSize, CondClass Cond, OrderedNode* Cmp1, OrderedNode* Cmp2) {
return _Select(OpSize::i64Bit, CompareSize, Cond, Cmp1, Cmp2, _InlineConstant(1), _InlineConstant(0));
}
IRPair<IROp_Select> To01(FEXCore::IR::OpSize CompareSize, OrderedNode* Cmp1) {
return Select01(CompareSize, CondClassType {COND_NEQ}, Cmp1, Constant(0));
return Select01(CompareSize, CondClass::NEQ, Cmp1, Constant(0));
}
IRPair<IROp_NZCVSelect> _NZCVSelect01(CondClassType Cond) {
IRPair<IROp_NZCVSelect> _NZCVSelect01(CondClass Cond) {
return _NZCVSelect(OpSize::i64Bit, Cond, _InlineConstant(1), _InlineConstant(0));
}
@@ -250,7 +300,7 @@ public:
}
/** @} */
FEXCore::IR::RegisterClassType WalkFindRegClass(OrderedNodeWrapper ssa) {
RegClass WalkFindRegClass(OrderedNodeWrapper ssa) {
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
return WalkFindRegClass(RealNode);
}
+1 -1
View File
@@ -39,7 +39,7 @@ public:
}
protected:
PassManager* Manager;
PassManager* Manager {};
};
class PassManager final {
@@ -93,11 +93,11 @@ void IRValidation::Run(IREmitter* IREmit) {
// After RA, the destination needs to be assigned a register and class
auto PhyReg = PhysicalRegister(CodeNode);
FEXCore::IR::RegisterClassType ExpectedClass = IR::GetRegClass(IROp->Op);
FEXCore::IR::RegisterClassType AssignedClass = FEXCore::IR::RegisterClassType {PhyReg.Class};
const auto ExpectedClass = IR::GetRegClass(IROp->Op);
const auto AssignedClass = PhyReg.AsRegClass();
// If no register class was assigned
if (AssignedClass == IR::InvalidClass) {
if (AssignedClass == IR::RegClass::Invalid) {
HadError |= true;
Errors << "%" << ID << ": Had destination but with no register class assigned" << std::endl;
}
@@ -109,10 +109,10 @@ void IRValidation::Run(IREmitter* IREmit) {
}
// Assigned class wasn't the expected class and it is a non-complex op
if (AssignedClass != ExpectedClass && ExpectedClass != IR::ComplexClass) {
if (AssignedClass != ExpectedClass && ExpectedClass != IR::RegClass::Complex) {
HadWarning |= true;
Warnings << "%" << ID << ": Destination had register class " << AssignedClass.Val << " When register class "
<< ExpectedClass.Val << " Was expected" << std::endl;
Warnings << "%" << ID << ": Destination had register class " << uint32_t(AssignedClass) << " When register class "
<< uint32_t(ExpectedClass) << " Was expected" << std::endl;
}
}
}
@@ -2,22 +2,21 @@
/*
$info$
tags: ir|opts
desc: This is not used right now, possibly broken
$end_info$
*/
#include "FEXCore/Core/X86Enums.h"
#include "FEXCore/Utils/CompilerDefs.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/fextl/deque.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/Profiler.h>
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/fextl/deque.h>
#include <FEXCore/fextl/vector.h>
// Flag bit flags
#define FLAG_V (1U << 0)
#define FLAG_C (1U << 1)
@@ -62,36 +61,36 @@ struct FlagInfo {
return {.Raw = R};
}
bool Trivial() {
bool Trivial() const {
return Raw == 0;
}
unsigned Read() {
unsigned Read() const {
return Bits(0, 8);
}
unsigned Write() {
unsigned Write() const {
return Bits(8, 8);
}
bool CanEliminate() {
bool CanEliminate() const {
return Bits(16, 1);
}
bool Special() {
bool Special() const {
return Bits(63, 1);
}
IROps Replacement() {
IROps Replacement() const {
return (IROps)Bits(32, 16);
}
IROps ReplacementNoWrite() {
IROps ReplacementNoWrite() const {
return (IROps)Bits(48, 16);
}
private:
unsigned Bits(unsigned Start, unsigned Count) {
unsigned Bits(unsigned Start, unsigned Count) const {
return (Raw >> Start) & ((1u << Count) - 1);
}
};
@@ -154,45 +153,44 @@ public:
private:
FlagInfo Classify(IROp_Header* Node);
unsigned FlagForReg(unsigned Reg);
unsigned FlagsForCondClassType(CondClassType Cond);
unsigned FlagsForCondClassType(CondClass Cond);
bool EliminateDeadCode(IREmitter* IREmit, Ref CodeNode, IROp_Header* IROp);
void FoldBranch(IREmitter* IREmit, IRListView& CurrentIR, IROp_CondJump* Op, Ref CodeNode);
CondClassType X86ToArmFloatCond(CondClassType X86);
CondClass X86ToArmFloatCond(CondClass X86);
bool ProcessBlock(IREmitter* IREmit, IRListView& CurrentIR, Ref Block, ControlFlowGraph& CFG);
void OptimizeParity(IREmitter* IREmit, IRListView& CurrentIR, ControlFlowGraph& CFG);
};
unsigned DeadFlagCalculationEliminination::FlagsForCondClassType(CondClassType Cond) {
unsigned DeadFlagCalculationEliminination::FlagsForCondClassType(CondClass Cond) {
switch (Cond) {
case COND_AL: return 0;
case CondClass::AL: return 0;
case COND_MI:
case COND_PL: return FLAG_N;
case CondClass::MI:
case CondClass::PL: return FLAG_N;
case COND_EQ:
case COND_NEQ: return FLAG_Z;
case CondClass::EQ:
case CondClass::NEQ: return FLAG_Z;
case COND_UGE:
case COND_ULT: return FLAG_C;
case CondClass::UGE:
case CondClass::ULT: return FLAG_C;
case COND_VS:
case COND_VC:
case COND_FU:
case COND_FNU: return FLAG_V;
case CondClass::VS:
case CondClass::VC:
case CondClass::FU:
case CondClass::FNU: return FLAG_V;
case COND_UGT:
case COND_ULE: return FLAG_Z | FLAG_C;
case CondClass::UGT:
case CondClass::ULE: return FLAG_Z | FLAG_C;
case COND_SGE:
case COND_SLT:
case COND_FLU:
case COND_FGE: return FLAG_N | FLAG_V;
case CondClass::SGE:
case CondClass::SLT:
case CondClass::FLU:
case CondClass::FGE: return FLAG_N | FLAG_V;
case COND_SGT:
case COND_SLE:
case COND_FLEU:
case COND_FGT: return FLAG_N | FLAG_Z | FLAG_V;
case CondClass::SGT:
case CondClass::SLE:
case CondClass::FLEU:
case CondClass::FGT: return FLAG_N | FLAG_Z | FLAG_V;
default: LOGMAN_THROW_A_FMT(false, "unknown cond class type"); return FLAG_NZCV;
}
@@ -456,7 +454,7 @@ bool DeadFlagCalculationEliminination::EliminateDeadCode(IREmitter* IREmit, Ref
return true;
}
CondClassType DeadFlagCalculationEliminination::X86ToArmFloatCond(CondClassType X86) {
CondClass DeadFlagCalculationEliminination::X86ToArmFloatCond(CondClass X86) {
// Table of x86 condition codes that map to arm64 condition codes, in the
// sense that fcmp+axflag+branch(x86) is equivalent to fcmp+branch(arm).
//
@@ -465,12 +463,12 @@ CondClassType DeadFlagCalculationEliminination::X86ToArmFloatCond(CondClassType
//
// SF/OF conditions are trivial and therefore shouldn't actually be generated
switch (X86) {
case COND_UGE /* A */: return {COND_FGE} /* GE */;
case COND_UGT /* AE */: return {COND_FGT} /* GT */;
case COND_ULT /* B */: return {COND_SLT} /* LT */;
case COND_ULE /* BE */: return {COND_SLE} /* LE */;
case COND_SLE /* LE */: return {COND_SLE} /* LE */;
default: return {COND_AL};
case CondClass::UGE /* A */: return CondClass::FGE /* GE */;
case CondClass::UGT /* AE */: return CondClass::FGT /* GT */;
case CondClass::ULT /* B */: return CondClass::SLT /* LT */;
case CondClass::ULE /* BE */: return CondClass::SLE /* LE */;
case CondClass::SLE /* LE */: return CondClass::SLE /* LE */;
default: return CondClass::AL;
}
}
@@ -485,8 +483,8 @@ void DeadFlagCalculationEliminination::FoldBranch(IREmitter* IREmit, IRListView&
auto Prev = CurrentIR.GetOp<IR::IROp_Header>(PrevWrap);
if (Prev->Op == OP_AXFLAG) {
// Pattern match a branch fed by AXFLAG.
CondClassType ArmCond = X86ToArmFloatCond(Op->Cond);
if (ArmCond == COND_AL) {
CondClass ArmCond = X86ToArmFloatCond(Op->Cond);
if (ArmCond == CondClass::AL) {
return;
}
@@ -495,7 +493,7 @@ void DeadFlagCalculationEliminination::FoldBranch(IREmitter* IREmit, IRListView&
// Pattern match a branch fed by a compare. We could also handle bit tests
// here, but tbz/tbnz has a limited offset range which we don't have a way to
// deal with yet. Let's hope that's not a big deal.
if (!(Op->Cond == COND_NEQ || Op->Cond == COND_EQ) || (Prev->Size < OpSize::i32Bit)) {
if (!(Op->Cond == CondClass::NEQ || Op->Cond == CondClass::EQ) || (Prev->Size < OpSize::i32Bit)) {
return;
}
@@ -629,9 +627,9 @@ void DeadFlagCalculationEliminination::OptimizeParity(IREmitter* IREmit, IRListV
}
for (auto [Block, BlockHeader] : CurrentIR.GetBlocks()) {
auto ID = BlockHeader->C<IROp_CodeBlock>()->ID;
const auto ID = BlockHeader->C<IROp_CodeBlock>()->ID;
const auto& Predecessors = CFG.Get(ID)->Predecessors;
bool Full = false;
auto Predecessors = CFG.Get(ID)->Predecessors;
if (Predecessors.empty()) {
// Conservatively assume there was full parity before the start block
@@ -12,6 +12,7 @@ $end_info$
#include "Interface/IR/Passes.h"
#include "Interface/Core/CPUID.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/fextl/vector.h>
@@ -22,7 +23,7 @@ using namespace FEXCore;
namespace FEXCore::IR {
namespace {
struct RegisterClass {
struct RegisterClassData {
uint32_t Available;
uint32_t Count;
@@ -32,9 +33,9 @@ namespace {
Ref RegToSSA[32];
};
IR::RegisterClassType GetRegClassFromNode(IR::IRListView* IR, IR::IROp_Header* IROp) {
IR::RegisterClassType Class = IR::GetRegClass(IROp->Op);
if (Class != IR::ComplexClass) {
IR::RegClass GetRegClassFromNode(IR::IRListView* IR, IR::IROp_Header* IROp) {
const auto Class = IR::GetRegClass(IROp->Op);
if (Class != IR::RegClass::Complex) {
return Class;
}
@@ -46,7 +47,7 @@ namespace {
case IR::OP_LOADMEM:
case IR::OP_LOADMEMTSO: return IROp->C<IR::IROp_LoadMem>()->Class;
case IR::OP_FILLREGISTER: return IROp->C<IR::IROp_FillRegister>()->Class;
default: return IR::InvalidClass;
default: return IR::RegClass::Invalid;
}
};
} // Anonymous namespace
@@ -56,15 +57,15 @@ public:
explicit ConstrainedRAPass(const FEXCore::CPUIDEmu* CPUID)
: CPUID {CPUID} {}
void Run(IREmitter* IREmit) override;
void AddRegisters(IR::RegisterClassType Class, uint32_t RegisterCount) override;
void AddRegisters(IR::RegClass Class, uint32_t RegisterCount) override;
bool TryPostRAMerge(Ref LastNode, Ref CodeNode, IROp_Header* IROp);
private:
RegisterClass Classes[IR::NumClasses];
RegisterClassData Classes[IR::NumClasses];
IREmitter* IREmit;
IRListView* IR;
const FEXCore::CPUIDEmu* CPUID;
IREmitter* IREmit {};
IRListView* IR {};
const FEXCore::CPUIDEmu* CPUID {};
// Map of nodes to their preferred register, to coalesce load/store reg.
fextl::vector<PhysicalRegister> PreferredReg;
@@ -82,7 +83,7 @@ private:
fextl::vector<bool> Seen;
// SourcesNextUses is read backwards, this tracks the index
int64_t SourceIndex;
int64_t SourceIndex {};
bool Rematerializable(IROp_Header* IROp) {
return IROp->Op == OP_CONSTANT;
@@ -101,7 +102,7 @@ private:
uint32_t SlotPlusOne = SpillSlots[IR->GetID(Node).Value];
LOGMAN_THROW_A_FMT(SlotPlusOne >= 1, "Node must have been spilled");
RegisterClassType RegClass = GetRegClassFromNode(IR, IROp);
const auto RegClass = GetRegClassFromNode(IR, IROp);
return IREmit->_FillRegister(IROp->Size, IROp->ElementSize, SlotPlusOne - 1, RegClass);
};
@@ -109,7 +110,7 @@ private:
// block, so we don't need to size the block up-front.
fextl::vector<uint32_t> NextUses;
bool AnySpilled;
bool AnySpilled {};
bool IsValidArg(OrderedNodeWrapper Arg) {
if (Arg.IsInvalid()) {
@@ -120,7 +121,7 @@ private:
return Op != OP_INLINECONSTANT && Op != OP_INLINEENTRYPOINTOFFSET;
};
RegisterClass* GetClass(PhysicalRegister Reg) {
RegisterClassData* GetClass(PhysicalRegister Reg) {
return &Classes[Reg.Class];
};
@@ -133,13 +134,13 @@ private:
LOGMAN_THROW_A_FMT(ID < SSAToReg.size(), "Only old nodes looked up");
PhysicalRegister Reg = SSAToReg[ID];
RegisterClass* Class = GetClass(Reg);
RegisterClassData* Class = GetClass(Reg);
return (Class->Available & GetRegBits(Reg)) == 0 && Class->RegToSSA[Reg.Reg] == Node;
};
void FreeReg(PhysicalRegister Reg) {
RegisterClass* Class = GetClass(Reg);
RegisterClassData* Class = GetClass(Reg);
uint32_t RegBits = GetRegBits(Reg);
LOGMAN_THROW_A_FMT(!(Class->Available & RegBits), "Register double-free");
@@ -187,22 +188,22 @@ private:
};
PhysicalRegister DecodeSRAReg(const IROp_Header* IROp, Ref Node) {
uint8_t FlagOffset = Classes[GPRFixedClass.Val].Count - 2;
uint8_t FlagOffset = Classes[FEXCore::ToUnderlying(RegClass::GPRFixed)].Count - 2;
if (IROp->Op == OP_STOREREGISTER) {
return PhysicalRegister(Node);
} else if (IROp->Op == OP_LOADPF || IROp->Op == OP_STOREPF) {
return PhysicalRegister {GPRFixedClass, FlagOffset};
return PhysicalRegister {RegClass::GPRFixed, FlagOffset};
} else if (IROp->Op == OP_LOADAF || IROp->Op == OP_STOREAF) {
return PhysicalRegister {GPRFixedClass, (uint8_t)(FlagOffset + 1)};
return PhysicalRegister {RegClass::GPRFixed, uint8_t(FlagOffset + 1)};
} else {
const IROp_LoadRegister* Op = IROp->C<IR::IROp_LoadRegister>();
LOGMAN_THROW_A_FMT(Op->Class == GPRClass || Op->Class == FPRClass, "SRA classes");
if (Op->Class == FPRClass) {
return PhysicalRegister {FPRFixedClass, (uint8_t)Op->Reg};
LOGMAN_THROW_A_FMT(Op->Class == RegClass::GPR || Op->Class == RegClass::FPR, "SRA classes");
if (Op->Class == RegClass::FPR) {
return PhysicalRegister {RegClass::FPRFixed, uint8_t(Op->Reg)};
} else {
return PhysicalRegister {GPRFixedClass, (uint8_t)Op->Reg};
return PhysicalRegister {RegClass::GPRFixed, uint8_t(Op->Reg)};
}
}
};
@@ -267,7 +268,7 @@ private:
SourceIndex = SourcesNextUses.size();
}
void SpillReg(RegisterClass* Class, IROp_CodeBlock* Block, IROp_Header* Exclude) {
void SpillReg(RegisterClassData* Class, IROp_CodeBlock* Block, IROp_Header* Exclude) {
// We're about to use next-use information, so calculate it.
if (!AnySpilled) {
CalculateNextUses(Block, Exclude);
@@ -318,7 +319,7 @@ private:
// If we already spilled the Candidate, we don't need to spill again.
// Similarly, if we can rematerialize the instruction, we don't spill it.
if (!Spilled && Header->Op != OP_CONSTANT) {
LOGMAN_THROW_A_FMT(Reg.Class == GetRegClassFromNode(IR, Header), "Consistent");
LOGMAN_THROW_A_FMT(Reg.AsRegClass() == GetRegClassFromNode(IR, Header), "Consistent");
// SpillSlots allocation is deferred.
if (SpillSlots.empty()) {
@@ -329,7 +330,7 @@ private:
uint32_t Slot = IR->GetHeader()->SpillSlots++;
// We must map here in case we're spilling something we shuffled.
auto SpillOp = IREmit->_SpillRegister(OrderedNodeWrapper::FromImmediate(Reg.Raw), Slot, RegisterClassType {Reg.Class});
auto SpillOp = IREmit->_SpillRegister(OrderedNodeWrapper::FromImmediate(Reg.Raw), Slot, Reg.AsRegClass());
SpillOp.first->Header.Size = Header->Size;
SpillOp.first->Header.ElementSize = Header->ElementSize;
SpillSlots[Value] = Slot + 1;
@@ -341,7 +342,7 @@ private:
};
void RemapReg(Ref Node, PhysicalRegister Reg) {
RegisterClass* Class = GetClass(Reg);
RegisterClassData* Class = GetClass(Reg);
Class->RegToSSA[Reg.Reg] = Node;
uint32_t Index = IR->GetID(Node).Value;
@@ -352,7 +353,7 @@ private:
// Record a given assignment of register Reg to Node.
void SetReg(Ref Node, PhysicalRegister Reg) {
RegisterClass* Class = GetClass(Reg);
RegisterClassData* Class = GetClass(Reg);
uint32_t RegBits = GetRegBits(Reg);
LOGMAN_THROW_A_FMT((Class->Available & RegBits) == RegBits, "Precondition");
@@ -370,7 +371,7 @@ private:
// Prioritize preferred registers.
if (Node < PreferredReg.size()) {
if (PhysicalRegister Reg = PreferredReg[Node]; !Reg.IsInvalid()) {
RegisterClass* Class = GetClass(Reg);
RegisterClassData* Class = GetClass(Reg);
uint32_t RegBits = GetRegBits(Reg);
if ((Class->Available & RegBits) == RegBits) {
@@ -383,10 +384,10 @@ private:
// Try to handle tied registers. This can fail, the JIT will insert moves.
if (int TiedIdx = IR::TiedSource(IROp->Op); TiedIdx >= 0) {
auto Reg = PhysicalRegister(IROp->Args[TiedIdx]);
RegisterClass* Class = GetClass(Reg);
RegisterClassData* Class = GetClass(Reg);
uint32_t RegBits = GetRegBits(Reg);
if (Reg.Class != GPRFixedClass && Reg.Class != FPRFixedClass && (Class->Available & RegBits) == RegBits) {
if (Reg.AsRegClass() != RegClass::GPRFixed && Reg.AsRegClass() != RegClass::FPRFixed && (Class->Available & RegBits) == RegBits) {
SetReg(CodeNode, Reg);
return;
}
@@ -394,7 +395,7 @@ private:
// Try to coalesce reserved pairs. Just a heuristic to remove some moves.
if (IROp->Op == OP_ALLOCATEGPR && IROp->C<IROp_AllocateGPR>()->ForPair) {
uint32_t Available = Classes[GPRClass].Available;
uint32_t Available = Classes[FEXCore::ToUnderlying(RegClass::GPR)].Available;
// Only choose base register R if R and R + 1 are both free
Available &= (Available >> 1);
@@ -405,20 +406,20 @@ private:
if (Available) {
unsigned Reg = std::countr_zero(Available);
SetReg(CodeNode, PhysicalRegister(GPRClass, Reg));
SetReg(CodeNode, PhysicalRegister(RegClass::GPR, Reg));
return;
}
} else if (IROp->Op == OP_ALLOCATEGPRAFTER) {
uint32_t Available = Classes[GPRClass].Available;
uint32_t Available = Classes[FEXCore::ToUnderlying(RegClass::GPR)].Available;
auto After = PhysicalRegister(IROp->Args[0]);
if ((After.Reg & 1) == 0 && Available & (1ull << (After.Reg + 1))) {
SetReg(CodeNode, PhysicalRegister(GPRClass, After.Reg + 1));
SetReg(CodeNode, PhysicalRegister(RegClass::GPR, After.Reg + 1));
return;
}
}
RegisterClassType ClassType = GetRegClassFromNode(IR, IROp);
RegisterClass* Class = &Classes[ClassType];
RegClass ClassType = GetRegClassFromNode(IR, IROp);
RegisterClassData* Class = &Classes[FEXCore::ToUnderlying(ClassType)];
// Spill to make room in the register file.
if (!Class->Available) {
@@ -433,10 +434,10 @@ private:
};
};
void ConstrainedRAPass::AddRegisters(IR::RegisterClassType Class, uint32_t RegisterCount) {
void ConstrainedRAPass::AddRegisters(IR::RegClass Class, uint32_t RegisterCount) {
LOGMAN_THROW_A_FMT(RegisterCount <= 31, "Up to 31 regs supported");
Classes[Class].Count = RegisterCount;
Classes[FEXCore::ToUnderlying(Class)].Count = RegisterCount;
}
inline bool KillMove(IROp_Header* LastOp, IROp_Header* IROp, Ref LastNode, Ref CodeNode) {
@@ -530,7 +531,7 @@ bool ConstrainedRAPass::TryPostRAMerge(Ref LastNode, Ref CodeNode, IROp_Header*
const auto Result = CPUID->RunFunction(ConstantFunction, 0 /* leaf */);
IREmit->SetWriteCursorBefore(CodeNode);
IREmit->_Fence({FEXCore::IR::Fence_Inst});
IREmit->_Fence(IR::FenceType::Inst);
IREmit->_Constant(Result.eax).Node->Reg = PhysicalRegister(Op->OutEAX).Raw;
IREmit->_Constant(Result.ebx).Node->Reg = PhysicalRegister(Op->OutEBX).Raw;
IREmit->_Constant(Result.ecx).Node->Reg = PhysicalRegister(Op->OutECX).Raw;
@@ -663,7 +664,7 @@ void ConstrainedRAPass::Run(IREmitter* IREmit_) {
// Static registers must be consistent at SRA load/store. Evict to ensure.
if (auto Node = DecodeSRANode(IROp, CodeNode); Node != nullptr) {
auto Reg = DecodeSRAReg(IROp, CodeNode);
RegisterClass* Class = &Classes[Reg.Class];
RegisterClassData* Class = &Classes[Reg.Class];
if (!(Class->Available & (1u << Reg.Reg))) {
Ref Old = Class->RegToSSA[Reg.Reg];
@@ -678,7 +679,7 @@ void ConstrainedRAPass::Run(IREmitter* IREmit_) {
Ref Copy;
if (Reg.Class == FPRFixedClass) {
if (Reg.AsRegClass() == RegClass::FPRFixed) {
IROp_Header* Header = IR->GetOp<IROp_Header>(Old);
Copy = IREmit->_VMov(Header->Size, OrderedNodeWrapper::FromImmediate(Reg.Raw));
} else {
@@ -12,14 +12,14 @@ $end_info$
#include <stdint.h>
namespace FEXCore::IR {
struct RegisterClassType;
enum class RegClass : uint32_t;
class RegisterAllocationPass : public FEXCore::IR::Pass {
public:
virtual void AddRegisters(FEXCore::IR::RegisterClassType Class, uint32_t RegisterCount) = 0;
virtual void AddRegisters(RegClass Class, uint32_t RegisterCount) = 0;
// Number of GPRs usable for pairs at start of GPR set. Must be even.
uint32_t PairRegs;
uint32_t PairRegs {};
};
} // namespace FEXCore::IR
@@ -6,7 +6,6 @@
#include "Interface/IR/PassManager.h"
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/Profiler.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/Core/HostFeatures.h"
#include "Interface/Core/Addressing.h"
@@ -66,7 +65,7 @@ public:
int8_t TopOffset = 0;
FixedSizeStack()
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T()}) {}
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T::Invalid}) {}
void push(const T& Value) {
rotate();
@@ -85,7 +84,7 @@ public:
}
void pop() {
buffer.front() = {StackSlot::INVALID, T()};
buffer.front() = {StackSlot::INVALID, T::Invalid};
rotate(false);
}
@@ -103,7 +102,7 @@ public:
void clear() {
for (auto& Elem : buffer) {
Elem = {StackSlot::UNUSED, T()};
Elem = {StackSlot::UNUSED, T::Invalid};
}
TopOffset = 0;
}
@@ -171,28 +170,40 @@ private:
// Helpers
Ref RotateRight8(uint32_t V, Ref Amount);
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
IREmit->_StoreMem(FPRClass, OpSize::i64Bit, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
IREmit->_StoreMemFPR(OpSize::i64Bit, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
auto Upper = IREmit->_VExtractToGPR(OpSize::i128Bit, OpSize::i64Bit, StackNode, 1);
// Store the Upper part of the register (the remaining 2 bytes) into memory.
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MEM_OFFSET_SXTX,
.IndexScale = OffsetScale,
.Offset = 8,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
A = SelectAddressMode(IREmit, A, GPROpSize, Features.SupportsTSOImm9, false, false, OpSize::i16Bit);
IREmit->_StoreMem(GPRClass, OpSize::i16Bit, Upper, A.Base, A.Index, OpSize::i64Bit, MEM_OFFSET_SXTX, A.IndexScale);
IREmit->_StoreMemGPR(OpSize::i16Bit, Upper, A.Base, A.Index, OpSize::i64Bit, MemOffsetType::SXTX, A.IndexScale);
}
void Store80BitToMem(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else {
F80SplitStore_Helper(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
}
void StoreStackMem_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(!ReducedPrecisionMode, "Full precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -204,22 +215,12 @@ private:
case OpSize::i32Bit:
case OpSize::i64Bit: {
StackNode = IREmit->_F80CVT(Op->StoreSize, StackNode);
IREmit->_StoreMem(FPRClass, Op->StoreSize, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
IREmit->_StoreMemFPR(Op->StoreSize, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
case OpSize::f80Bit: {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MEM_OFFSET_SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else { // 80bit requires split-store
F80SplitStore_Helper(Op, StackNode);
}
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -229,6 +230,8 @@ private:
// Performs a store to memory from a value the stack passed in as StackNode.
// This is the version dealing with the reduced precision case.
void StoreStackMem_Reduced_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(ReducedPrecisionMode, "Reduced precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -241,14 +244,13 @@ private:
[[fallthrough]];
}
case OpSize::i64Bit: {
IREmit->_StoreMem(FPRClass, Op->StoreSize, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
IREmit->_StoreMemFPR(Op->StoreSize, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
// 80bit requires split-store
case OpSize::f80Bit: {
StackNode = IREmit->_F80CVTTo(StackNode, OpSize::i64Bit);
F80SplitStore_Helper(Op, StackNode);
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -290,23 +292,24 @@ private:
void Reset();
struct StackMemberInfo {
StackMemberInfo() {}
StackMemberInfo() = delete;
StackMemberInfo(Ref Data)
: StackDataNode(Data) {}
StackMemberInfo(Ref Data, Ref Source, OpSize Size, bool Float)
StackMemberInfo(Ref Data, Ref Source, OpSize Size)
: StackDataNode(Data)
, Source({Size, Source})
, InterpretAsFloat(Float) {}
, Source({Size, Source}) {}
Ref StackDataNode {}; // Reference to the data in the Stack.
// This is the source data node in the stack format, possibly converted to 64/80 bits.
struct StackMemberData final {
OpSize Size;
Ref Node;
};
static const StackMemberInfo Invalid;
// Tuple is only valid if we have information about the Source of the Stack Data Node.
// In it's valid then OpSize is the original source size and Ref is the original source node.
std::optional<StackMemberData> Source {};
bool InterpretAsFloat {false}; // True if this is a floating point value, false if integer
};
// StackData, TopCache need to be always properly set to ensure
@@ -359,6 +362,8 @@ private:
IRListView* IR = nullptr;
};
inline const X87StackOptimization::StackMemberInfo X87StackOptimization::StackMemberInfo::Invalid {nullptr};
inline void X87StackOptimization::InvalidateCaches() {
InvalidateCachedRegs();
ConstantPool.fill(nullptr);
@@ -401,8 +406,7 @@ inline void X87StackOptimization::MigrateToSlowPathIf(bool ShouldMigrate) {
inline Ref X87StackOptimization::GetTopWithCache_Slow() {
if (!TopOffsetCache[0]) {
TopOffsetCache[0] =
IREmit->_LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
TopOffsetCache[0] = IREmit->_LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
return TopOffsetCache[0];
}
@@ -447,7 +451,7 @@ inline void X87StackOptimization::SetTopWithCache_Slow(Ref Value) {
inline Ref X87StackOptimization::GetFTW() {
if (!FTWCached) {
FTWCached = IREmit->_LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
FTWCached = IREmit->_LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
}
return FTWCached;
}
@@ -470,7 +474,7 @@ inline Ref X87StackOptimization::LoadStackValueAtOffset_Slow(uint8_t Offset) {
OrderedNode* TopOffsetAddress = GetOffsetTopAddressWithCache_Slow(Offset);
auto Size = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
if (!TopValueCache[Offset]) {
TopValueCache[Offset] = IREmit->_LoadMem(FPRClass, Size, TopOffsetAddress, IREmit->_InlineConstant(MMBaseOffset()), Size, MEM_OFFSET_SXTX, 1);
TopValueCache[Offset] = IREmit->_LoadMemFPR(Size, TopOffsetAddress, IREmit->_InlineConstant(MMBaseOffset()), Size, MemOffsetType::SXTX, 1);
}
return TopValueCache[Offset];
}
@@ -595,7 +599,7 @@ inline void X87StackOptimization::UpdateTopForPush_Slow() {
void X87StackOptimization::FlushCachedRegs() {
if (FlushTopPending) {
IREmit->_StoreContext(OpSize::i8Bit, GPRClass, TopOffsetCache[0], offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
IREmit->_StoreContextGPR(OpSize::i8Bit, TopOffsetCache[0], offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
FlushTopPending = false;
}
@@ -603,7 +607,7 @@ void X87StackOptimization::FlushCachedRegs() {
for (size_t i = 0; i < FlushValuesPending.size(); i++) {
if (FlushValuesPending[i]) {
OrderedNode* TopOffsetAddress = GetOffsetTopAddressWithCache_Slow(i);
IREmit->_StoreMem(FPRClass, Size, TopValueCache[i], TopOffsetAddress, IREmit->_InlineConstant(MMBaseOffset()), Size, MEM_OFFSET_SXTX, 1);
IREmit->_StoreMemFPR(Size, TopValueCache[i], TopOffsetAddress, IREmit->_InlineConstant(MMBaseOffset()), Size, MemOffsetType::SXTX, 1);
// store
FlushValuesPending[i] = false;
}
@@ -654,7 +658,7 @@ void X87StackOptimization::FlushCachedRegs() {
}
}();
IREmit->_StoreContext(OpSize::i8Bit, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
IREmit->_StoreContextGPR(OpSize::i8Bit, NewFTW, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
FTWCached = NewFTW;
}
@@ -729,6 +733,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// The optimization should run per-block
Reset();
IREmit->SetCurrentCodeBlock(BlockNode);
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (!LoweredX87(IROp->Op)) {
continue;
@@ -928,8 +933,13 @@ void X87StackOptimization::Run(IREmitter* Emit) {
StoreStackValueAtOffset_Slow(SourceNode);
} else {
auto* SourceNode = CurrentIR.GetNode(Op->X80Src);
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize, Op->Float});
if (Op->OriginalValue.IsInvalid()) {
// No original value to track - just push the converted data
StackData.push(StackMemberInfo {SourceNode});
} else {
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize});
}
}
break;
}
@@ -994,9 +1004,16 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// str w2, [x1]
// or similar. As long as the source size and dest size are one and the same.
// This will avoid any conversions between source and stack element size and conversion back.
if (!SlowPath && Value->Source && Value->Source->Size == Op->StoreSize && Value->InterpretAsFloat) {
IREmit->_StoreMem(Value->InterpretAsFloat ? FPRClass : GPRClass, Op->StoreSize, Value->Source->Node, AddrNode, Offset, Align,
OffsetType, OffsetScale);
OpSize StoreSize = Op->StoreSize;
LOGMAN_THROW_A_FMT(Op->StoreSize == OpSize::i32Bit || Op->StoreSize == OpSize::i64Bit || Op->StoreSize == OpSize::f80Bit,
"Invalid store size in x87 store stack mem");
if (!SlowPath && Value->Source && Value->Source->Size == StoreSize) {
Ref SourceValue = Value->Source->Node;
if (Op->StoreSize == OpSize::f80Bit) {
Store80BitToMem(Op, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
} else {
IREmit->_StoreMemFPR(StoreSize, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
break;
}
@@ -1036,11 +1053,26 @@ void X87StackOptimization::Run(IREmitter* Emit) {
case OP_F80STACKXCHANGE: {
const auto* Op = IROp->C<IROp_F80StackXchange>();
auto Offset = Op->SrcStack;
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
if (Offset == 0) {
// No-op
break;
}
const auto [ValidTop, StackMemberTop] = StackData.top(0);
const auto [ValidOffset, StackMemberOffset] = StackData.top(Offset);
if (ValidTop != StackSlot::VALID || ValidOffset != StackSlot::VALID) {
// Slow path: do actual memory operations
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
} else {
// Fast path: swap complete StackMemberInfo preserving Source metadata
StackData.setTop(StackMemberOffset, 0);
StackData.setTop(StackMemberTop, Offset);
}
break;
}
@@ -1160,7 +1192,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
Ref Value {};
if (ReducedPrecisionMode) {
Value = IREmit->_Vector_FToI(OpSize::i64Bit, OpSize::i64Bit, St0, Round_Host);
Value = IREmit->_Vector_FToI(OpSize::i64Bit, OpSize::i64Bit, St0, RoundMode::Host);
} else {
Value = IREmit->_F80Round(St0);
}
@@ -19,9 +19,9 @@ union PhysicalRegister {
return Raw == Other.Raw;
}
PhysicalRegister(RegisterClassType Class, uint8_t Reg)
PhysicalRegister(RegClass Class, uint8_t Reg)
: Reg(Reg)
, Class(Class.Val) {}
, Class(uint8_t(Class)) {}
PhysicalRegister(OrderedNodeWrapper Arg)
: Raw(Arg.GetImmediate()) {}
@@ -29,12 +29,16 @@ union PhysicalRegister {
PhysicalRegister(Ref Node)
: Raw(Node->Reg) {}
RegClass AsRegClass() const {
return RegClass {Class};
}
static const PhysicalRegister Invalid() {
return PhysicalRegister(InvalidClass, 0);
return PhysicalRegister(RegClass::Invalid, 0);
}
bool IsInvalid() const {
static_assert(InvalidClass == 0);
static_assert(uint8_t(RegClass::Invalid) == 0);
return Raw == 0;
}
};
+14 -7
View File
@@ -4,6 +4,7 @@
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/PrctlUtils.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/memory.h>
@@ -50,8 +51,17 @@ void* FEX_mmap(void* addr, size_t length, int prot, int flags, int fd, off_t off
errno = -(uint64_t)Result;
return (void*)-1;
}
if (flags & MAP_ANONYMOUS) {
VirtualName("FEXMem", Result, length);
}
return Result;
}
void VirtualName(const char* Name, void* Ptr, size_t Size) {
prctl(PR_SET_VMA, PR_SET_VMA_ANON_NAME, Ptr, Size, Name);
}
int FEX_munmap(void* addr, size_t length) {
int Result = Alloc64->Munmap(addr, length);
@@ -83,7 +93,7 @@ void* DisableSBRKAllocations() {
// calls won't allocate any memory through that.
void* AlignedBRK = reinterpret_cast<void*>(FEXCore::AlignUp(reinterpret_cast<uintptr_t>(StartingSBRK), FEXCore::Utils::FEX_PAGE_SIZE));
void* AfterBRK =
mmap(AlignedBRK, FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_FIXED_NOREPLACE | MAP_NORESERVE, -1, 0);
::mmap(AlignedBRK, FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_FIXED_NOREPLACE | MAP_NORESERVE, -1, 0);
if (AfterBRK == INVALID_PTR) {
// Couldn't allocate the page after the aligned brk? This should never happen.
// FEXCore::LogMan isn't configured yet so we just need to print the message.
@@ -130,10 +140,7 @@ void ClearHooks() {
FEXCore::Allocator::mmap = ::mmap;
FEXCore::Allocator::munmap = ::munmap;
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Alloc64.release();
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Alloc64));
}
#pragma GCC diagnostic pop
@@ -284,7 +291,7 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
--StackRegionIt;
auto Alloc =
mmap(StackRegionIt->Ptr, StackRegionIt->Size, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED, -1, 0);
::mmap(StackRegionIt->Ptr, StackRegionIt->Size, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED, -1, 0);
LogMan::Throw::AFmt(Alloc != MAP_FAILED, "mmap({},{:x}) failed", fmt::ptr(StackRegionIt->Ptr), StackRegionIt->Size);
LogMan::Throw::AFmt(Alloc == StackRegionIt->Ptr, "mmap returned {} instead of {}", Alloc, fmt::ptr(StackRegionIt->Ptr));
@@ -295,7 +302,7 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
// Block remaining memory gaps
for (auto RegionIt = Regions.begin(); RegionIt != Regions.end(); ++RegionIt) {
auto Alloc = mmap(RegionIt->Ptr, RegionIt->Size, PROT_NONE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0);
auto Alloc = ::mmap(RegionIt->Ptr, RegionIt->Size, PROT_NONE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0);
LogMan::Throw::AFmt(Alloc != MAP_FAILED, "mmap({},{:x}) failed", fmt::ptr(RegionIt->Ptr), RegionIt->Size);
LogMan::Throw::AFmt(Alloc == RegionIt->Ptr, "mmap returned {} instead of {}", Alloc, fmt::ptr(RegionIt->Ptr));
@@ -98,7 +98,7 @@ private:
// Align UsedPages so it pads to the next page.
// Necessary to take advantage of madvise zero page pooling.
using FlexBitElementType = uint64_t;
alignas(4096) FEXCore::FlexBitSet<FlexBitElementType> UsedPages;
alignas(FEXCore::Utils::FEX_PAGE_SIZE) FEXCore::FlexBitSet<FlexBitElementType> UsedPages;
// This returns the size of the LiveVMARegion in addition to the flex set that tracks the used data
// The LiveVMARegion lives at the start of the VMA region which means on initialization we need to set that
@@ -140,7 +140,7 @@ private:
}
};
static_assert(sizeof(LiveVMARegion) == 4096, "Needs to be the size of a page");
static_assert(sizeof(LiveVMARegion) == FEXCore::Utils::FEX_PAGE_SIZE, "Needs to be the size of a page");
static_assert(std::is_trivially_copyable<LiveVMARegion>::value, "Needs to be trivially copyable");
static_assert(offsetof(LiveVMARegion, UsedPages) == sizeof(LiveVMARegion), "FlexBitSet needs to be at the end");
@@ -168,6 +168,7 @@ private:
LOGMAN_THROW_A_FMT(Res != -1, "Couldn't mprotect region: {} '{}' Likely occurs when running out of memory or Maximum VMAs", errno,
strerror(errno));
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(ReservedRegion->Base), SizePlusManagedData);
LiveVMARegion* LiveRange = new (reinterpret_cast<void*>(ReservedRegion->Base)) LiveVMARegion();
// Copy over the reserved data
@@ -206,7 +207,7 @@ OSAllocator_64Bit::LiveVMARegion* OSAllocator_64Bit::FindLiveRegionForAddress(ui
uintptr_t RegionBegin = (*it)->SlabInfo->Base;
uintptr_t RegionEnd = RegionBegin + (*it)->SlabInfo->RegionSize;
if (Addr >= RegionBegin && Addr < RegionEnd) {
if (Addr >= RegionBegin && AddrEnd < RegionEnd) {
LiveRegion = *it;
// Leave our loop
break;
@@ -404,14 +405,18 @@ again:
// Mark the pages as used
uintptr_t RegionBegin = LiveRegion->SlabInfo->Base;
uintptr_t MappedBegin = (AllocatedOffset - RegionBegin) >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t PagesSet {};
for (size_t i = 0; i < NumberOfPages; ++i) {
LiveRegion->UsedPages.Set(MappedBegin + i);
PagesSet += LiveRegion->UsedPages.TestAndSet(MappedBegin + i) == false;
}
// Change our last allocation region
LiveRegion->LastPageAllocation = MappedBegin + NumberOfPages;
LiveRegion->FreeSpace -= length;
LiveRegion->FreeSpace -= PagesSet * FEXCore::Utils::FEX_PAGE_SIZE;
LOGMAN_THROW_A_FMT(LiveRegion->FreeSpace <= LiveRegion->SlabInfo->RegionSize,
"Corrupt LiveRegion free space! 0x{:x} > 0x{:x}. After allocating 0x{:x} (0x{:x} overlapped)", LiveRegion->FreeSpace,
LiveRegion->SlabInfo->RegionSize, length, PagesSet);
}
if (!AllocatedOffset) {
@@ -473,7 +478,7 @@ int OSAllocator_64Bit::Munmap(void* addr, size_t length) {
::mmap(addr, length, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_FIXED, -1, 0);
}
(*it)->FreeSpace += FreedPages * 4096;
(*it)->FreeSpace += FreedPages * FEXCore::Utils::FEX_PAGE_SIZE;
// Set the last allocated page to the minimum of last page allocation or this slab
// This will let us more quickly fill holes
@@ -505,6 +510,8 @@ void OSAllocator_64Bit::AllocateMemoryRegions(fextl::vector<FEXCore::Allocator::
// This enables the kernel to use transparent large pages in the allocator which can reduce memory pressure
::madvise(it.Ptr, ObjectAllocSize, MADV_HUGEPAGE);
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(it.Ptr), ObjectAllocSize);
ObjectAlloc = new (it.Ptr) Alloc::ForwardOnlyIntrusiveArenaAllocator(it.Ptr, ObjectAllocSize);
ReservedRegions = ObjectAlloc->new_construct(ReservedRegions, ObjectAlloc);
LiveRegions = ObjectAlloc->new_construct(LiveRegions, ObjectAlloc);
@@ -602,6 +609,8 @@ fextl::unique_ptr<T> make_alloc_unique(FEXCore::Allocator::MemoryRegion& Base, A
ERROR_AND_DIE_FMT("Couldn't allocate memory region");
}
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(ptr), MinPage);
// Remove the page from the base region.
// Could be zero after this.
Base.Size -= MinPage;
+20 -6
View File
@@ -27,6 +27,11 @@ struct FlexBitSet final {
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
return Value;
}
bool TestAndSet(size_t Element) {
bool Value = Get(Element);
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
return Value;
}
void Set(size_t Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
@@ -70,12 +75,17 @@ struct FlexBitSet final {
template<bool WantUnset>
BitsetScanResults BackwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t MinimumElement) {
bool FoundHole {};
for (size_t CurrentPage = BeginningElement; CurrentPage >= (MinimumElement + ElementCount);) {
// Final element to iterate to.
const size_t FinalElement = MinimumElement + ElementCount - 1;
for (size_t CurrentPage = BeginningElement; CurrentPage >= FinalElement;) {
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT(Remaining <= CurrentPage, "Scanning less than available range");
LOGMAN_THROW_A_FMT(CurrentPage <= BeginningElement && CurrentPage >= FinalElement, "BackwardScanForRange: Scanning less than "
"available range");
while (Remaining) {
if (this->Get(CurrentPage - Remaining) == WantUnset) {
if (this->Get(CurrentPage - Remaining + 1) == WantUnset) {
// Has an intersecting range
break;
}
@@ -92,7 +102,7 @@ struct FlexBitSet final {
CurrentPage -= Remaining;
} else {
// We have a slab range
return BitsetScanResults {CurrentPage - ElementCount, FoundHole};
return BitsetScanResults {CurrentPage - ElementCount + 1, FoundHole};
}
}
@@ -108,11 +118,15 @@ struct FlexBitSet final {
BitsetScanResults ForwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t ElementsInSet) {
bool FoundHole {};
for (size_t CurrentElement = BeginningElement; CurrentElement < (ElementsInSet - ElementCount);) {
// Final element to iterate to.
const size_t FinalElement = ElementsInSet - ElementCount + 1;
for (size_t CurrentElement = BeginningElement; CurrentElement <= FinalElement;) {
// If we have enough free space, check if we have enough free pages that are contiguous
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT((CurrentElement + Remaining - 1) < ElementsInSet, "Scanning less than available range");
LOGMAN_THROW_A_FMT(CurrentElement >= BeginningElement && CurrentElement <= FinalElement, "ForwardScanForRange: Scanning less than "
"available range");
while (Remaining) {
if (this->Get(CurrentElement + Remaining - 1) == WantUnset) {
@@ -53,4 +53,12 @@ public:
namespace Alloc::OSAllocator {
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocator();
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocatorWithRegions(fextl::vector<FEXCore::Allocator::MemoryRegion>& Regions);
static inline void ReleaseAllocatorWorkaround(fextl::unique_ptr<Alloc::HostAllocator> Allocator) {
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// The allocator is also intrusively allocated, so the unique_ptr tries to double free the HostAllocator object.
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Allocator.release();
}
} // namespace Alloc::OSAllocator
+102 -122
View File
@@ -144,33 +144,33 @@ static __uint128_t LoadAcquire128(uint64_t Addr) {
}
static uint64_t LoadAcquire64(uint64_t Addr) {
std::atomic<uint64_t>* Atom = reinterpret_cast<std::atomic<uint64_t>*>(Addr);
return Atom->load(std::memory_order_acquire);
auto Atom = std::atomic_ref<uint64_t>(*reinterpret_cast<uint64_t*>(Addr));
return Atom.load(std::memory_order_acquire);
}
static bool StoreCAS64(uint64_t& Expected, uint64_t Val, uint64_t Addr) {
std::atomic<uint64_t>* Atom = reinterpret_cast<std::atomic<uint64_t>*>(Addr);
return Atom->compare_exchange_strong(Expected, Val);
auto Atom = std::atomic_ref<uint64_t>(*reinterpret_cast<uint64_t*>(Addr));
return Atom.compare_exchange_strong(Expected, Val);
}
static uint32_t LoadAcquire32(uint64_t Addr) {
std::atomic<uint32_t>* Atom = reinterpret_cast<std::atomic<uint32_t>*>(Addr);
return Atom->load(std::memory_order_acquire);
auto Atom = std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(Addr));
return Atom.load(std::memory_order_acquire);
}
static bool StoreCAS32(uint32_t& Expected, uint32_t Val, uint64_t Addr) {
std::atomic<uint32_t>* Atom = reinterpret_cast<std::atomic<uint32_t>*>(Addr);
return Atom->compare_exchange_strong(Expected, Val);
auto Atom = std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(Addr));
return Atom.compare_exchange_strong(Expected, Val);
}
static uint8_t LoadAcquire8(uint64_t Addr) {
std::atomic<uint8_t>* Atom = reinterpret_cast<std::atomic<uint8_t>*>(Addr);
return Atom->load(std::memory_order_acquire);
auto Atom = std::atomic_ref<uint8_t>(*reinterpret_cast<uint8_t*>(Addr));
return Atom.load(std::memory_order_acquire);
}
static bool StoreCAS8(uint8_t& Expected, uint8_t Val, uint64_t Addr) {
std::atomic<uint8_t>* Atom = reinterpret_cast<std::atomic<uint8_t>*>(Addr);
return Atom->compare_exchange_strong(Expected, Val);
auto Atom = std::atomic_ref<uint8_t>(*reinterpret_cast<uint8_t*>(Addr));
return Atom.compare_exchange_strong(Expected, Val);
}
static uint16_t DoLoad16(uint64_t Addr) {
@@ -211,8 +211,8 @@ static uint16_t DoLoad16(uint64_t Addr) {
uint64_t Alignment = Addr & AlignmentMask;
Addr &= ~AlignmentMask;
std::atomic<uint64_t>* Atomic = reinterpret_cast<std::atomic<uint64_t>*>(Addr);
uint64_t TmpResult = Atomic->load();
auto Atomic = std::atomic_ref<uint64_t>(*reinterpret_cast<uint64_t*>(Addr));
uint64_t TmpResult = Atomic.load();
// Zexts the result
uint16_t Result = TmpResult >> (Alignment * 8);
@@ -224,8 +224,8 @@ static uint16_t DoLoad16(uint64_t Addr) {
uint64_t Alignment = Addr & AlignmentMask;
Addr &= ~AlignmentMask;
std::atomic<uint32_t>* Atomic = reinterpret_cast<std::atomic<uint32_t>*>(Addr);
uint32_t TmpResult = Atomic->load();
auto Atomic = std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(Addr));
uint32_t TmpResult = Atomic.load();
// Zexts the result
uint16_t Result = TmpResult >> (Alignment * 8);
@@ -272,8 +272,8 @@ static uint32_t DoLoad32(uint64_t Addr) {
uint64_t Alignment = Addr & AlignmentMask;
Addr &= ~AlignmentMask;
std::atomic<uint64_t>* Atomic = reinterpret_cast<std::atomic<uint64_t>*>(Addr);
uint64_t TmpResult = Atomic->load();
auto Atomic = std::atomic_ref<uint64_t>(*reinterpret_cast<uint64_t*>(Addr));
uint64_t TmpResult = Atomic.load();
return TmpResult >> (Alignment * 8);
}
@@ -465,7 +465,7 @@ static bool RunCASPAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg1, uint3
// Fits within a 16byte region
uint64_t Alignment = Addr & 0b1111;
Addr &= ~0b1111ULL;
std::atomic<__uint128_t>* Atomic128 = reinterpret_cast<std::atomic<__uint128_t>*>(Addr);
auto Atomic128 = std::atomic_ref<__uint128_t>(*reinterpret_cast<__uint128_t*>(Addr));
__uint128_t Mask = ~0ULL;
Mask <<= Alignment * 8;
@@ -480,7 +480,7 @@ static bool RunCASPAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg1, uint3
Expected <<= Alignment * 8;
while (1) {
TmpExpected = Atomic128->load();
TmpExpected = Atomic128.load();
// Set up expected
TmpExpected &= NegMask;
@@ -491,7 +491,7 @@ static bool RunCASPAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg1, uint3
TmpDesired &= NegMask;
TmpDesired |= Desired;
bool CASResult = Atomic128->compare_exchange_strong(TmpExpected, TmpDesired);
bool CASResult = Atomic128.compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Successful, so we are done
return true;
@@ -617,36 +617,6 @@ static uint64_t HandleCASPAL_ARMv8(uint32_t Instr, uintptr_t ProgramCounter, uin
}
}
static bool HandleAtomicVectorStore(uint32_t Instr, uintptr_t ProgramCounter) {
uint32_t* PC = (uint32_t*)ProgramCounter;
uint32_t Size = (Instr >> 30) & 1;
uint32_t DataReg = Instr & 0x1F;
if (Size == 1) {
// 64-bit pair happens on paranoid vector stores
// [0] ldaxp(xzr, TMP3, MemSrc); // <- Can hit SIGBUS. Overwritten with DMB
// [1] stlxp(TMP3, TMP1, TMP2, MemSrc); // <- Can also hit SIGBUS
// [2] cbnz(TMP3, &B); // < Overwritten with DMB
if (DataReg == 31) {
uint32_t NextInstr = PC[1];
uint32_t AddrReg = (NextInstr >> 5) & 0x1F;
DataReg = NextInstr & 0x1F;
uint32_t DataReg2 = (NextInstr >> 10) & 0x1F;
uint32_t STP = (0b10 << 30) | (0b101001000000000 << 15) | (DataReg2 << 10) | (AddrReg << 5) | DataReg;
PC[0] = DMB;
PC[1] = STP;
PC[2] = DMB;
// Back up one instruction and have another go
ClearICache(&PC[0], 12);
return true;
}
}
return false;
}
template<typename T>
using CASExpectedFn = T (*)(T Src, T Expected);
template<typename T>
@@ -740,7 +710,7 @@ static uint16_t DoCAS16(uint16_t DesiredSrc, uint16_t ExpectedSrc, uint64_t Addr
// Fits within a 16byte region
uint64_t Alignment = Addr & 0b1111;
Addr &= ~0b1111ULL;
std::atomic<__uint128_t>* Atomic128 = reinterpret_cast<std::atomic<__uint128_t>*>(Addr);
auto Atomic128 = std::atomic_ref<__uint128_t>(*reinterpret_cast<__uint128_t*>(Addr));
__uint128_t Mask = 0xFFFF;
Mask <<= Alignment * 8;
@@ -749,7 +719,7 @@ static uint16_t DoCAS16(uint16_t DesiredSrc, uint16_t ExpectedSrc, uint64_t Addr
__uint128_t TmpDesired {};
while (1) {
TmpExpected = Atomic128->load();
TmpExpected = Atomic128.load();
__uint128_t Desired = DesiredFunction(TmpExpected >> (Alignment * 8), DesiredSrc);
Desired <<= Alignment * 8;
@@ -766,7 +736,7 @@ static uint16_t DoCAS16(uint16_t DesiredSrc, uint16_t ExpectedSrc, uint64_t Addr
TmpDesired &= NegMask;
TmpDesired |= Desired;
bool CASResult = Atomic128->compare_exchange_strong(TmpExpected, TmpDesired);
bool CASResult = Atomic128.compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Successful, so we are done
return Expected >> (Alignment * 8);
@@ -810,9 +780,9 @@ static uint16_t DoCAS16(uint16_t DesiredSrc, uint16_t ExpectedSrc, uint64_t Addr
uint64_t TmpExpected {};
uint64_t TmpDesired {};
std::atomic<uint64_t>* Atomic = reinterpret_cast<std::atomic<uint64_t>*>(Addr);
auto Atomic = std::atomic_ref<uint64_t>(*reinterpret_cast<uint64_t*>(Addr));
while (1) {
TmpExpected = Atomic->load();
TmpExpected = Atomic.load();
uint64_t Desired = DesiredFunction(TmpExpected >> (Alignment * 8), DesiredSrc);
Desired <<= Alignment * 8;
@@ -829,7 +799,7 @@ static uint16_t DoCAS16(uint16_t DesiredSrc, uint16_t ExpectedSrc, uint64_t Addr
TmpDesired &= NegMask;
TmpDesired |= Desired;
bool CASResult = Atomic->compare_exchange_strong(TmpExpected, TmpDesired);
bool CASResult = Atomic.compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Successful, so we are done
return Expected >> (Alignment * 8);
@@ -873,9 +843,9 @@ static uint16_t DoCAS16(uint16_t DesiredSrc, uint16_t ExpectedSrc, uint64_t Addr
uint32_t TmpExpected {};
uint32_t TmpDesired {};
std::atomic<uint32_t>* Atomic = reinterpret_cast<std::atomic<uint32_t>*>(Addr);
auto Atomic = std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(Addr));
while (1) {
TmpExpected = Atomic->load();
TmpExpected = Atomic.load();
uint32_t Desired = DesiredFunction(TmpExpected >> (Alignment * 8), DesiredSrc);
@@ -893,7 +863,7 @@ static uint16_t DoCAS16(uint16_t DesiredSrc, uint16_t ExpectedSrc, uint64_t Addr
TmpDesired &= NegMask;
TmpDesired |= Desired;
bool CASResult = Atomic->compare_exchange_strong(TmpExpected, TmpDesired);
bool CASResult = Atomic.compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Successful, so we are done
return Expected >> (Alignment * 8);
@@ -1040,7 +1010,7 @@ static uint32_t DoCAS32(uint32_t DesiredSrc, uint32_t ExpectedSrc, uint64_t Addr
// Fits within a 16byte region
uint64_t Alignment = Addr & 0b1111;
Addr &= ~0b1111ULL;
std::atomic<__uint128_t>* Atomic128 = reinterpret_cast<std::atomic<__uint128_t>*>(Addr);
auto Atomic128 = std::atomic_ref<__uint128_t>(*reinterpret_cast<__uint128_t*>(Addr));
__uint128_t Mask = ~0U;
Mask <<= Alignment * 8;
@@ -1049,7 +1019,7 @@ static uint32_t DoCAS32(uint32_t DesiredSrc, uint32_t ExpectedSrc, uint64_t Addr
__uint128_t TmpDesired {};
while (1) {
__uint128_t TmpActual = Atomic128->load();
__uint128_t TmpActual = Atomic128.load();
__uint128_t Desired = DesiredFunction(TmpActual >> (Alignment * 8), DesiredSrc);
__uint128_t Expected = ExpectedFunction(TmpActual >> (Alignment * 8), ExpectedSrc);
@@ -1064,7 +1034,7 @@ static uint32_t DoCAS32(uint32_t DesiredSrc, uint32_t ExpectedSrc, uint64_t Addr
TmpDesired &= NegMask;
TmpDesired |= Desired << (Alignment * 8);
bool CASResult = Atomic128->compare_exchange_strong(TmpExpected, TmpDesired);
bool CASResult = Atomic128.compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Stored successfully
return Expected;
@@ -1108,9 +1078,9 @@ static uint32_t DoCAS32(uint32_t DesiredSrc, uint32_t ExpectedSrc, uint64_t Addr
uint64_t TmpExpected {};
uint64_t TmpDesired {};
std::atomic<uint64_t>* Atomic = reinterpret_cast<std::atomic<uint64_t>*>(Addr);
auto Atomic = std::atomic_ref<uint64_t>(*reinterpret_cast<uint64_t*>(Addr));
while (1) {
uint64_t TmpActual = Atomic->load();
uint64_t TmpActual = Atomic.load();
uint64_t Desired = DesiredFunction(TmpActual >> (Alignment * 8), DesiredSrc);
uint64_t Expected = ExpectedFunction(TmpActual >> (Alignment * 8), ExpectedSrc);
@@ -1125,7 +1095,7 @@ static uint32_t DoCAS32(uint32_t DesiredSrc, uint32_t ExpectedSrc, uint64_t Addr
TmpDesired &= NegMask;
TmpDesired |= Desired << (Alignment * 8);
bool CASResult = Atomic->compare_exchange_strong(TmpExpected, TmpDesired);
bool CASResult = Atomic.compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Stored successfully
return Expected;
@@ -1270,7 +1240,7 @@ static uint64_t DoCAS64(uint64_t DesiredSrc, uint64_t ExpectedSrc, uint64_t Addr
// Fits within a 16byte region
uint64_t Alignment = Addr & AlignmentMask;
Addr &= ~AlignmentMask;
std::atomic<__uint128_t>* Atomic128 = reinterpret_cast<std::atomic<__uint128_t>*>(Addr);
auto Atomic128 = std::atomic_ref<__uint128_t>(*reinterpret_cast<__uint128_t*>(Addr));
__uint128_t Mask = ~0ULL;
Mask <<= Alignment * 8;
@@ -1279,7 +1249,7 @@ static uint64_t DoCAS64(uint64_t DesiredSrc, uint64_t ExpectedSrc, uint64_t Addr
__uint128_t TmpDesired {};
while (1) {
__uint128_t TmpActual = Atomic128->load();
__uint128_t TmpActual = Atomic128.load();
__uint128_t Desired = DesiredFunction(TmpActual >> (Alignment * 8), DesiredSrc);
__uint128_t Expected = ExpectedFunction(TmpActual >> (Alignment * 8), ExpectedSrc);
@@ -1294,7 +1264,7 @@ static uint64_t DoCAS64(uint64_t DesiredSrc, uint64_t ExpectedSrc, uint64_t Addr
TmpDesired &= NegMask;
TmpDesired |= Desired << (Alignment * 8);
bool CASResult = Atomic128->compare_exchange_strong(TmpExpected, TmpDesired);
bool CASResult = Atomic128.compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Stored successfully
return Expected;
@@ -1326,9 +1296,7 @@ static uint64_t DoCAS64(uint64_t DesiredSrc, uint64_t ExpectedSrc, uint64_t Addr
}
}
static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_t ExpectedReg, uint32_t AddressReg, uint32_t* StrictSplitLockMutex) {
uint64_t Addr = GPRs[AddressReg];
static std::optional<uint64_t> DoCAS(uint32_t Size, uint64_t Desired, uint64_t Expected, uint64_t Addr, uint32_t* StrictSplitLockMutex) {
// Cross-cacheline CAS doesn't work on ARM
// It isn't even guaranteed to work on x86
// Intel will do a "split lock" which locks the full bus
@@ -1341,7 +1309,7 @@ static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_
// Only need to handle 16, 32, 64
if (Size == 2) {
auto Res = DoCAS16<false>(
GPRs[DesiredReg], GPRs[ExpectedReg], Addr,
Desired, Expected, Addr,
[](uint16_t, uint16_t Expected) -> uint16_t {
// Expected is just Expected
return Expected;
@@ -1351,16 +1319,10 @@ static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_
return Desired;
},
StrictSplitLockMutex);
// Regardless of pass or fail
// We set the result register if it isn't a zero register
if (ExpectedReg != 31) {
GPRs[ExpectedReg] = Res;
}
return true;
return Res;
} else if (Size == 4) {
auto Res = DoCAS32<false>(
GPRs[DesiredReg], GPRs[ExpectedReg], Addr,
Desired, Expected, Addr,
[](uint32_t, uint32_t Expected) -> uint32_t {
// Expected is just Expected
return Expected;
@@ -1370,16 +1332,10 @@ static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_
return Desired;
},
StrictSplitLockMutex);
// Regardless of pass or fail
// We set the result register if it isn't a zero register
if (ExpectedReg != 31) {
GPRs[ExpectedReg] = Res;
}
return true;
return Res;
} else if (Size == 8) {
auto Res = DoCAS64<false>(
GPRs[DesiredReg], GPRs[ExpectedReg], Addr,
Desired, Expected, Addr,
[](uint64_t, uint64_t Expected) -> uint64_t {
// Expected is just Expected
return Expected;
@@ -1389,16 +1345,24 @@ static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_
return Desired;
},
StrictSplitLockMutex);
// Regardless of pass or fail
// We set the result register if it isn't a zero register
if (ExpectedReg != 31) {
GPRs[ExpectedReg] = Res;
}
return true;
return Res;
}
return false;
return std::nullopt;
}
static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_t ExpectedReg, uint32_t AddressReg, uint32_t* StrictSplitLockMutex) {
std::optional<uint64_t> Res = DoCAS(Size, GPRs[DesiredReg], GPRs[ExpectedReg], GPRs[AddressReg], StrictSplitLockMutex);
if (!Res.has_value()) {
return false;
}
// Regardless of pass or fail
// We set the result register if it isn't a zero register
if (ExpectedReg != 31) {
GPRs[ExpectedReg] = *Res;
}
return true;
}
static bool HandleCASAL(uint64_t* GPRs, uint32_t Instr, uint32_t* StrictSplitLockMutex) {
@@ -1560,38 +1524,43 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
return false;
}
static bool HandleAtomicLoad(uint32_t Instr, uint64_t* GPRs, int64_t Offset) {
static bool HandleAtomicLoad(uint32_t Instr, uint64_t* GPRs, int64_t Offset, Core::UnalignedExclusiveStore* Store = nullptr) {
uint32_t Size = 1 << (Instr >> 30);
uint32_t ResultReg = Instr & 0b11111;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
uint64_t Addr = GPRs[AddressReg] + Offset;
uint64_t Res;
if (Size == 2) {
auto Res = DoLoad16(Addr);
Res = DoLoad16(Addr);
// We set the result register if it isn't a zero register
if (ResultReg != 31) {
GPRs[ResultReg] = Res;
}
return true;
} else if (Size == 4) {
auto Res = DoLoad32(Addr);
Res = DoLoad32(Addr);
// We set the result register if it isn't a zero register
if (ResultReg != 31) {
GPRs[ResultReg] = Res;
}
return true;
} else if (Size == 8) {
auto Res = DoLoad64(Addr);
Res = DoLoad64(Addr);
// We set the result register if it isn't a zero register
if (ResultReg != 31) {
GPRs[ResultReg] = Res;
}
return true;
} else {
return false;
}
return false;
if (Store) {
Store->Addr = Addr;
Store->Store = Res;
Store->Size = Size;
}
return true;
}
static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, uint32_t* StrictSplitLockMutex) {
@@ -1952,8 +1921,8 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
}
[[nodiscard]]
std::optional<int32_t>
HandleUnalignedAccess(FEXCore::Core::InternalThreadState* Thread, UnalignedHandlerType HandleType, uintptr_t ProgramCounter, uint64_t* GPRs) {
std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState* Thread, UnalignedHandlerType HandleType,
uintptr_t ProgramCounter, uint64_t* GPRs, bool IsJIT) {
#ifdef _M_ARM_64
constexpr bool is_arm64 = true;
#else
@@ -1977,8 +1946,7 @@ HandleUnalignedAccess(FEXCore::Core::InternalThreadState* Thread, UnalignedHandl
auto CTX = static_cast<Context::ContextImpl*>(Thread->CTX);
uint32_t* StrictSplitLockMutex {CTX->Config.StrictInProcessSplitLocks ? &CTX->StrictSplitLockMutex : nullptr};
// ParanoidTSO path doesn't modify any code.
if (HandleType == UnalignedHandlerType::Paranoid) [[unlikely]] {
if (!IsJIT) [[unlikely]] {
if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if (ArchHelpers::Arm64::HandleAtomicLoad(Instr, GPRs, 0)) {
@@ -2016,7 +1984,29 @@ HandleUnalignedAccess(FEXCore::Core::InternalThreadState* Thread, UnalignedHandl
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDLUR*: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & ArchHelpers::Arm64::LDAXR_MASK) == ArchHelpers::Arm64::LDAXR_INST) { // LDAXR*
if (ArchHelpers::Arm64::HandleAtomicLoad(Instr, GPRs, 0, &Thread->ExclusiveStore)) {
return 4;
}
} else if ((Instr & ArchHelpers::Arm64::STLXR_MASK) == ArchHelpers::Arm64::STLXR_INST) { // STLXR*
uint32_t StatusReg = Instr << 11 >> 27;
// // Emulate exclusive store by validating the address and value against the last unaligned LDAXR*.
if (GPRs[AddrReg] != Thread->ExclusiveStore.Addr || Size > Thread->ExclusiveStore.Size) {
if (StatusReg != 31) {
GPRs[StatusReg] = 1;
}
return 4;
}
if (std::optional<uint64_t> Prev =
DoCAS(Size, DataReg == 31 ? 0 : GPRs[DataReg], Thread->ExclusiveStore.Store, GPRs[AddrReg], StrictSplitLockMutex)) {
if (StatusReg != 31) {
GPRs[StatusReg] = !!memcmp(&Thread->ExclusiveStore.Store, &*Prev, Size);
}
Thread->ExclusiveStore.Size = 0;
return 4;
}
}
return 0;
}
const auto Frame = Thread->CurrentFrame;
@@ -2068,6 +2058,9 @@ HandleUnalignedAccess(FEXCore::Core::InternalThreadState* Thread, UnalignedHandl
if (BytesToSkip) {
// Skip this instruction now
return BytesToSkip;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASPAL: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
}
@@ -2131,19 +2124,6 @@ HandleUnalignedAccess(FEXCore::Core::InternalThreadState* Thread, UnalignedHandl
ClearICache(&PC[-1], 8);
// Back up one instruction and have another go
return -4;
} else if ((Instr & ArchHelpers::Arm64::LDAXP_MASK) == ArchHelpers::Arm64::LDAXP_INST) { // LDAXP
/// This is handling the case of paranoid ARMv8.0-a atomic stores.
/// This backpatches the ldaxp+stlxp+cbnz if the previous `HandleCASPAL_ARMv8` didn't handle the case.
if (ArchHelpers::Arm64::HandleAtomicVectorStore(Instr, ProgramCounter)) {
return 0;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAXP: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & ArchHelpers::Arm64::STLXP_MASK) == ArchHelpers::Arm64::STLXP_INST) { // STLXP
// Should not trigger - middle of an LDAXP/STAXP pair.
LogMan::Msg::EFmt("Unhandled JIT SIGBUS STLXP: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
// Check if another thread backpatched this instruction before this thread got here
+32 -4
View File
@@ -1,7 +1,10 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/Utils/LogManager.h>
namespace FEXCore::LongJump {
#include <cstring>
namespace FEXCore::UncheckedLongJump {
#if defined(_M_ARM_64)
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
@@ -32,7 +35,7 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value) {
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
// x0 contains the jumpbuffer
ldp x19, x20, [x0, #( 0 * 8)];
@@ -58,6 +61,27 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value)
)" ::
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
// First 12 values are registers [x19,x30].
memcpy(&GPRs[19], &Buffer.Registers[0], sizeof(uint64_t) * 12);
// Next 8 values are [D8,D15]
// Retain upper 64-bits of the register, only modifying lower 64-bits.
for (size_t i = 0; i < 8; ++i) {
memcpy(&FPRs[8 + i], &Buffer.Registers[12 + i], sizeof(uint64_t));
}
// Last value is stack pointer
memcpy(&GPRs[31], &Buffer.Registers[20], sizeof(uint64_t));
// Load the expected value in to X0
GPRs[0] = Value;
// Load the PC with the current LR.
*PC = GPRs[30];
}
#else
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
@@ -86,7 +110,7 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value) {
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
.intel_syntax noprefix;
// rdi contains the jumpbuffer
@@ -115,5 +139,9 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value)
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
LOGMAN_MSG_A_FMT("This is unimplemented on x86-64");
}
#endif
} // namespace FEXCore::LongJump
} // namespace FEXCore::UncheckedLongJump
+5
View File
@@ -9,6 +9,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/string.h>
@@ -70,6 +71,10 @@ static std::array<const char*, 2> TraceFSDirectories {
};
void Init() {
FEX_CONFIG_OPT(EnableGpuvisProfiling, ENABLEGPUVISPROFILING);
if (!EnableGpuvisProfiling()) {
return;
}
for (auto Path : TraceFSDirectories) {
#ifdef _WIN32
constexpr auto flags = O_WRONLY;
+62 -40
View File
@@ -1,9 +1,14 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <chrono>
#include <mutex>
#include <type_traits>
#include <FEXCore/fextl/functional.h>
#include <FEXCore/Utils/EnumUtils.h>
namespace FEXCore::Utils::SpinWaitLock {
/**
* @brief This provides routines to implement implement an "efficient spin-loop" using ARM's WFE and exclusive monitor interfaces.
@@ -123,35 +128,26 @@ static inline uint64_t WFELoadAtomic(uint64_t* Futex) {
return Result;
}
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
std::atomic<T>* AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Result = AtomicFutex->load();
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
return;
}
do {
while (!Pred {}(Result, ComparisonValue)) {
Result = LoadExclusive(Futex);
if (Result == ExpectedValue) {
if (Pred {}(Result, ComparisonValue)) {
return;
}
Result = WFELoadAtomic(Futex);
} while (Result != ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
Result = WFELoadAtomic(Futex);
}
}
template<typename T, typename TT>
static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanoseconds& Timeout) {
std::atomic<T>* AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex->load();
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
@@ -184,27 +180,43 @@ template bool Wait<uint16_t>(uint16_t*, uint16_t, const std::chrono::nanoseconds
template bool Wait<uint32_t>(uint32_t*, uint32_t, const std::chrono::nanoseconds&);
template bool Wait<uint64_t>(uint64_t*, uint64_t, const std::chrono::nanoseconds&);
#else
template<typename T, typename TT>
static inline void Wait(T* Futex, TT ExpectedValue) {
std::atomic<T>* AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Result = AtomicFutex->load();
template<typename T>
static inline T OneShotWFEBitComparison(T* Futex, T Mask, T Comp) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
return;
if ((Result & Mask) == Comp) {
return Result;
}
do {
Result = AtomicFutex->load();
} while (Result != ExpectedValue);
Result = LoadExclusive(Futex);
if ((Result & Mask) == Comp) {
return Result;
}
// Waits for write and returns result.
Result = WFELoadAtomic(Futex);
return Result;
}
#else
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
while (!Pred {}(Result, ComparisonValue)) {
Result = AtomicFutex.load();
}
}
template<typename T, typename TT>
static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanoseconds& Timeout) {
std::atomic<T>* AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex->load();
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
@@ -214,7 +226,7 @@ static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanosecon
const auto Begin = std::chrono::high_resolution_clock::now();
do {
Result = AtomicFutex->load();
Result = AtomicFutex.load();
const auto CurrentCycleCounter = std::chrono::high_resolution_clock::now();
if ((CurrentCycleCounter - Begin) >= Timeout) {
@@ -228,14 +240,24 @@ static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanosecon
}
#endif
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
WaitPred<std::equal_to<>, T>(Futex, ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
template<typename T>
static inline void lock(T* Futex) {
std::atomic<T>* AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Expected {};
T Desired {1};
// Try to CAS immediately.
if (AtomicFutex->compare_exchange_strong(Expected, Desired)) {
if (AtomicFutex.compare_exchange_strong(Expected, Desired)) {
return;
}
@@ -243,17 +265,17 @@ static inline void lock(T* Futex) {
// Wait until the futex is unlocked.
Wait(Futex, 0);
Expected = 0;
} while (!AtomicFutex->compare_exchange_strong(Expected, Desired));
} while (!AtomicFutex.compare_exchange_strong(Expected, Desired));
}
template<typename T>
static inline bool try_lock(T* Futex) {
std::atomic<T>* AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Expected {};
T Desired {1};
// Try to CAS immediately.
if (AtomicFutex->compare_exchange_strong(Expected, Desired)) {
if (AtomicFutex.compare_exchange_strong(Expected, Desired)) {
return true;
}
@@ -262,8 +284,8 @@ static inline bool try_lock(T* Futex) {
template<typename T>
static inline void unlock(T* Futex) {
std::atomic<T>* AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
AtomicFutex->store(0);
auto AtomicFutex = std::atomic_ref<T>(*Futex);
AtomicFutex.store(0);
}
#undef SPINLOOP_8BIT
+381
View File
@@ -0,0 +1,381 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <cstdint>
#if !defined(_WIN32)
#include <linux/futex.h> /* Definition of FUTEX_* constants */
#include <sys/syscall.h> /* Definition of SYS_* constants */
#include <unistd.h>
#else
#include <synchapi.h>
#endif
#include <FEXCore/Utils/LogManager.h>
#include "Utils/SpinWaitLock.h"
namespace FEXCore::Utils::WritePriorityMutex {
// A custom mutex that prioritizes exclusive locks.
// In highly contested scenarios, this can help minimize overall contention time.
//
// Features:
// - Up to 32767 pending exclusive locks ("writers")
// - Up to 32767 pending shared_locks ("readers")
// - Low-overhead waiting via WFE with a fallback to futex on timeout
// - Direct writer->reader hand-off and vice-versa to further reduce overhead
//
// Trade-offs:
// - No guaranteed order of wake-ups besides prioritizing writers
// - No support for recursive locking
// - We can't use FUTEX_LOCK_PI to enable priority inheritance
class Mutex final {
public:
Mutex() = default;
// Move-only type
Mutex(const Mutex&) = delete;
Mutex& operator=(const Mutex&) = delete;
Mutex(Mutex&& rhs) = delete;
Mutex& operator=(Mutex&&) = delete;
void lock() {
// Try a non-blocking lock first.
if (try_lock()) {
return;
}
// Try a quick WFE write-lock.
if (Attempt_WFE_WriteLock()) {
return;
}
// Still couldn't get it. Start waiting.
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected {};
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
// Increment the number of write waiters.
Desired = Expected + WRITE_WAITER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & WRITE_WAITER_COUNT_MASK) != 0, "Overflow in write-waiters!");
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Expected = AtomicFutex.fetch_add(WRITE_WAITER_INCREMENT);
Desired = Expected + WRITE_WAITER_INCREMENT;
#endif
// Thread added to waiter list.
Expected = Desired;
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & READ_OWNER_COUNT_MASK) == 0) {
// If not write-owned, and no read-owners, try to acquire.
LOGMAN_THROW_A_FMT((Expected & WRITE_WAITER_COUNT_MASK) != 0, "Underflow in write-waiters!");
// Add write-owned bit.
Desired = Expected | WRITE_OWNED_BIT;
// Remove ourselves from the wait list.
Desired -= WRITE_WAITER_INCREMENT;
Sleep = false;
} else {
// Already write-owned or read-locked. Go to sleep.
Desired = Expected;
Sleep = true;
break;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Somehow acquired a write-lock without it being set!");
return;
}
FutexWaitForWriteAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void lock_shared() {
// Try an uncontended lock first.
if (try_lock_shared()) {
return;
}
// Try a quick WFE read-lock.
if (Attempt_WFE_ReadLock()) {
return;
}
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
Sleep = false;
} else {
// Waiting for lock to become available. Add to waiters.
Desired = Expected | READ_WAITER_BIT;
Sleep = true;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Somehow read-locked and got a write lock!");
return;
}
FutexWaitForReadAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void unlock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Trying to write-unlock something not write-locked!");
// Remove the exclusive lock bit.
Desired = Expected & ~WRITE_OWNED_BIT;
// If no more writers, then make sure to clear the read-waiters bit as well.
if ((Desired & WRITE_WAITER_COUNT_MASK) == 0) {
Desired &= ~READ_WAITER_BIT;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
// If success, then `Expected` has old value. Containing `READ_WAITER_BIT` which was just masked off, and also `WRITE_WAITER_COUNT_MASK`.
if ((Expected & WRITE_WAITER_COUNT_MASK)) {
// Handle write-write handoff.
FutexWakeWriter();
} else if ((Expected & READ_WAITER_BIT)) {
// Handle write-reader handoff.
FutexWakeReaders();
}
}
void unlock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Trying to read-unlock something write-locked!");
LOGMAN_THROW_A_FMT((Expected & READ_OWNER_COUNT_MASK) != 0, "Trying to read-unlock something not read-locked!");
// Decrement the shared counter.
Desired = Expected - READ_OWNER_INCREMENT;
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Desired = AtomicFutex.fetch_sub(READ_OWNER_INCREMENT) - READ_OWNER_INCREMENT;
#endif
// Handle read->write handoff if there are any waiting writers, and no readers left.
if ((Desired & WRITE_WAITER_COUNT_MASK) && (Desired & READ_OWNER_COUNT_MASK) == 0) {
FutexWakeWriter();
}
}
bool try_lock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = 0;
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
// try to CAS immediately.
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
// Can race with other threads trying to lock shared!
bool try_lock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
// Exclusively owned or has a list of waiting owners. Can't pass.
if ((Expected & WRITE_OWNED_BIT) || (Expected & WRITE_WAITER_COUNT_MASK)) {
return false;
}
// Try to add reader.
uint32_t Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
// Uncontended mutex check
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
#if !defined(_WIN32)
// Initialize the internal mutex object to its default initializer state.
// Should only ever be used in the child process when a Linux fork() has occured.
void StealAndDropActiveLocks() {
Futex = 0;
}
#endif
private:
#if !defined(_WIN32)
void FutexWaitForWriteAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Read-lock waiting for writers to drain out.
void FutexWaitForReadAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
// Read-Lock or Write-lock unlocked, wake one writer.
// - Read->Write handoff.
// - Write->Write handoff.
void FutexWakeWriter() {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, 1, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Write-lock unlocked, wake read-locks waiting.
void FutexWakeReaders() {
// Wake all readers.
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, INT_MAX, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
#else
// Writers wait for the full 32-bit futex.
void FutexWaitForWriteAvailable(uint32_t Expected) {
WaitOnAddress(&Futex, &Expected, sizeof(Futex), INFINITE);
}
// Readers wait for Futex bits [31:16] to be zero.
void FutexWaitForReadAvailable(uint32_t Expected) {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
uint16_t smol_Expected = Expected >> 16;
WaitOnAddress(ReadWaiterAddress, &smol_Expected, sizeof(smol_Expected), INFINITE);
}
void FutexWakeWriter() {
WakeByAddressSingle(&Futex);
}
void FutexWakeReaders() {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
WakeByAddressAll(ReadWaiterAddress);
}
#endif
// Reuse the SpinWaitLock WFE implementations for read/write lock acquiring with WFE.
// Can't reuse the spin-lock directly as some bit-representations are different.
// WFE-write-lock is less likely to occur the more read-lock threads are participating. Can still occur so good to try.
// WFE-read-lock is actually quite likely to succeed.
// Return: true if the lock was acquired.
bool Attempt_WFE_WriteLock() {
#ifdef _M_ARM_64
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
while ((Now - Begin) < Duration) {
if (Expected == 0) {
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, ~0U, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
// Return: true if the lock was acquired.
bool Attempt_WFE_ReadLock() {
#ifdef _M_ARM_64
// Spin on a WFE for a short-amount of time, waiting for write-owned and writer-count to be zero.
// - Attempt to acquire read-lock at that point.
// - Don't add read-waiters bit on failure, return false.
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while ((Now - Begin) < Duration) {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, WRITE_OWNED_BIT | WRITE_WAITER_COUNT_MASK, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
constexpr static uint32_t WRITE_OWNED_BIT = 1U << 31;
constexpr static uint32_t READ_WAITER_BIT = 1U << 15;
constexpr static uint32_t WRITE_WAITER_OFFSET = 16;
constexpr static uint32_t WRITE_WAITER_INCREMENT = 1U << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_INCREMENT = 1;
// Count masks
constexpr static uint32_t WRITE_WAITER_COUNT_MASK = 0x7FFFU << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_COUNT_MASK = 0x7FFFU;
// Independent futex bit-set masks.
// Wait for readers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_READERS = 1U << 0;
// Wait for writers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_WRITERS = 1U << 1;
// Only spin on WFE for 0.01ms (10k ns).
constexpr static uint64_t CYCLECOUNT_DIVISOR = 1'000'000'000ULL / 10'000U;
// Layout:
// Bits[31]: Write-lock bit.
// Bits[30:16]: Write-waiter count.
// Bits[15]: Read-waiter bit.
// Bits[14:0]: Read-owner count.
uint32_t Futex {};
};
} // namespace FEXCore::Utils::WritePriorityMutex
+63 -33
View File
@@ -103,28 +103,25 @@ static inline std::optional<fextl::string> EnumParser(const ArrayPairType& EnumP
return fextl::fmt::format("{}", EnumMask);
}
namespace DefaultValues {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) extern const P(type) P(enum);
#define OPT_STR(group, enum, json, default) extern const std::string_view P(enum);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#include <FEXCore/Config/ConfigValues.inl>
using StringArrayType = fextl::list<fextl::string>;
namespace Type {
using StringArrayType = fextl::list<fextl::string>;
#define OPT_BASE(type, group, enum, json, default) using P(enum) = P(type);
#define OPT_STR(group, enum, json, default) using P(enum) = fextl::string;
#define OPT_STRARRAY(group, enum, json, default) using P(enum) = StringArrayType;
namespace detail {
template<ConfigOption Option>
struct ConfigOptionInfo;
#define DEFINE_METAINFO(type, enum, default) \
template<> \
struct ConfigOptionInfo<ConfigOption::CONFIG_##enum> { \
using Type = type; \
static auto Default() { \
extern default; \
return enum; \
} \
};
#define OPT_BASE(type, group, enum, json, default) DEFINE_METAINFO(type, enum, const type enum)
#define OPT_STR(group, enum, json, default) DEFINE_METAINFO(fextl::string, enum, const std::string_view enum)
#define OPT_STRARRAY(group, enum, json, default) DEFINE_METAINFO(StringArrayType, enum, const std::string_view enum)
#include <FEXCore/Config/ConfigValues.inl>
} // namespace Type
#define FEX_CONFIG_OPT(name, enum) \
FEXCore::Config::Value<FEXCore::Config::DefaultValues::Type::enum> name { \
FEXCore::Config::CONFIG_##enum, \
FEXCore::Config::DefaultValues::enum \
}
#undef P
} // namespace DefaultValues
} // namespace detail
FEX_DEFAULT_VISIBILITY void SetDataDirectory(std::string_view Path, bool Global);
FEX_DEFAULT_VISIBILITY void SetConfigDirectory(const std::string_view Path, bool Global);
@@ -135,8 +132,7 @@ FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigDirectory(bool Global);
FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigFileLocation(bool Global = false);
FEX_DEFAULT_VISIBILITY fextl::string GetApplicationConfig(const std::string_view Program, bool Global);
using LayerValue =
std::variant< fextl::string, DefaultValues::Type::StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerValue = std::variant< fextl::string, StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerOptions = fextl::unordered_map<ConfigOption, LayerValue>;
@@ -151,16 +147,16 @@ public:
return OptionMap.find(Option) != OptionMap.end();
}
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
std::optional<StringArrayType*> All(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
return std::nullopt;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
return &std::get<DefaultValues::Type::StringArrayType>(Value);
return &std::get<StringArrayType>(Value);
}
std::optional<fextl::string*> Get(ConfigOption Option) {
@@ -201,12 +197,12 @@ public:
auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
// If the option didn't exist as a StringArrayType yet, emplace it.
it = OptionMap.emplace(Option, DefaultValues::Type::StringArrayType {}).first;
it = OptionMap.emplace(Option, StringArrayType {}).first;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<DefaultValues::Type::StringArrayType>(Value).emplace_back(Data);
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<StringArrayType>(Value).emplace_back(Data);
}
void Erase(ConfigOption Option) {
@@ -236,7 +232,9 @@ FEX_DEFAULT_VISIBILITY fextl::string FindContainerPrefix();
FEX_DEFAULT_VISIBILITY void AddLayer(fextl::unique_ptr<FEXCore::Config::Layer> _Layer);
FEX_DEFAULT_VISIBILITY bool Exists(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<StringArrayType*> All(ConfigOption Option);
template<typename T>
FEX_DEFAULT_VISIBILITY std::optional<T> GetConv(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<fextl::string*> Get(ConfigOption Option);
FEX_DEFAULT_VISIBILITY void Set(ConfigOption Option, std::string_view Data);
FEX_DEFAULT_VISIBILITY void Erase(ConfigOption Option);
@@ -271,18 +269,18 @@ public:
return ValueData;
}
Value(T Value) requires (!std::is_same_v<T, DefaultValues::Type::StringArrayType>)
Value(T Value) requires (!std::is_same_v<T, StringArrayType>)
{
ValueData = std::move(Value);
}
// Array value types.
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, StringArrayType>)
{
GetListIfExists(Option, &ValueData);
}
DefaultValues::Type::StringArrayType& All() requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
StringArrayType& All() requires (std::is_same_v<T, StringArrayType>)
{
return ValueData;
}
@@ -293,6 +291,38 @@ private:
static T GetIfExists(FEXCore::Config::ConfigOption Option, T Default);
static T GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
};
/**
* Wrapper around Value that automatically picks the default for the given ConfigOption
*/
template<ConfigOption Option>
struct FEX_DEFAULT_VISIBILITY Getter : public Value<typename detail::ConfigOptionInfo<Option>::Type> {
using OptionInfo = detail::ConfigOptionInfo<Option>;
Getter()
: Value<typename OptionInfo::Type> {Option, OptionInfo::Default()} {}
};
/**
* Helper for reading a config value with caching.
*
* Typically this is used to declare class members so that the value is read
* on construction of the parent.
*/
#define FEX_CONFIG_OPT(name, enum) FEXCore::Config::Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> name {}
#define OPT_BASE(type, group, enum, json, default) \
/** \
* Helper for reading a config value. \
* \
* In contrast to FEX_CONFIG_OPT, this can be used in arbitrary expressions, \
* at the expense of not caching the value. Use Getter instead if the value \
* is read frequently. \
*/ \
inline auto Get_##enum() { \
return Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> {}; \
}
#include <FEXCore/Config/ConfigValues.inl>
} // namespace FEXCore::Config
+136 -1
View File
@@ -1,10 +1,20 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/functional.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <atomic>
#include <cstdint>
#include <mutex>
#include <optional>
#include <shared_mutex>
#include <span>
#include <unistd.h>
namespace FEXCore {
@@ -20,8 +30,14 @@ namespace HLE {
struct ExecutableFileInfo {
~ExecutableFileInfo();
#if __clang_major__ < 16
// Workaround for broken aggregate-initialization with std::piecewise_construct
ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap>, uint64_t, fextl::string);
ExecutableFileInfo() = default;
#endif
fextl::unique_ptr<HLE::SourcecodeMap> SourcecodeMap;
fextl::string FileId;
uint64_t FileId = 0;
fextl::string Filename;
};
@@ -33,10 +49,129 @@ struct ExecutableFileSectionInfo {
uintptr_t FileStartVA;
};
using CodeMapFileId = uint64_t;
/**
* Code maps capture information required for offline code cache generation
* and are written to disk during execution of FEX.
*
* Almost all CodeMap data will be an Entry that indicates blocks to be
* compiled for cache generation. The reserved value `LoadExternalLibrary`
* indicates that an instance of ExternalLibraryInfo follows (the entry data
* itself should be skipped in that case).
*/
struct CodeMap {
// Describes the location of an entry block compiled during execution
struct FEX_PACKED Entry {
CodeMapFileId FileId;
uint32_t BlockOffset;
};
// Describes an external library referenced during execution
struct ExternalLibraryInfo {
CodeMapFileId ExternalFileId;
// null-terminated file path; EITHER relative to the main executable OR an absolute path OR starting with a magic identifier:
// - WINE/: Path to Wine/Proton installation
// - WINEPREFIX/: Path to Wine/Proton prefix
// - SLR/: Path to Steam Linux Runtime
// At runtime, FEX will always dump absolute paths
char Path[];
// Followed by padding to a 4 byte boundary
};
// Followed by ExternalLibraryInfo
static constexpr Entry LoadExternalLibrary = {0xffff'ffff'ffff'ffff, 0xffff'ffff};
struct FEX_PACKED SetExecutableFileId {
Entry Marker = {0xffff'ffff'ffff'ffff, 0xffff'fffe};
CodeMapFileId ExecutableFileId;
};
struct ParsedContents {
fextl::string Filename;
fextl::set<uint64_t> Blocks;
bool IsExecutable = false;
};
// Follows scheme fileid[-nomb]
// The nomb ("no multiblock") suffix signifies that the code map is for use without multiblock, only.
static fextl::string GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix);
static fextl::map<CodeMapFileId, ParsedContents> ParseCodeMap(std::ifstream& File);
};
struct CodeMapOpener {
virtual ~CodeMapOpener() = default;
virtual int OpenCodeMapFile() = 0;
};
class CodeMapWriter {
public:
CodeMapWriter(CodeMapOpener&, bool OpenEagerly = false);
~CodeMapWriter();
// Checks if writing is enabled. Calls to this functions may also be interpreted as signals that writes are about to happen
bool IsWriteEnabled(const ExecutableFileSectionInfo&);
void ResetAfterFork() {
if (CodeMapFD.value_or(-1) != -1) {
close(CodeMapFD.value());
CodeMapFD.reset();
}
BufferOffset = 0;
KnownFileIds.clear();
}
bool IsBackingFD(int FD) const {
if (FD == CodeMapFD) {
LogMan::Msg::DFmt("Hiding directory entry for code map FD");
return true;
}
return false;
}
void AppendBlock(const FEXCore::ExecutableFileSectionInfo&, uint64_t Entry);
void AppendLibraryLoad(const FEXCore::ExecutableFileInfo&);
void AppendSetMainExecutable(const FEXCore::ExecutableFileInfo&);
// Thread-safely commit any pending data to disk
void Flush(size_t Offset);
private:
// Queues data into an internal ring buffer.
// Call Flush() to commit the data to disk.
void AppendData(std::span<const std::byte> Data);
// Commit given data range to disk
void Flush(size_t Offset, std::unique_lock<std::shared_mutex>&);
std::shared_mutex Mutex;
fextl::vector<std::byte> Buffer;
std::atomic<size_t> BufferOffset {0};
fextl::set<CodeMapFileId> KnownFileIds;
// std::nullopt: We haven't requested a CodeMapFD yet
// value is -1: We requested a CodeMapFD but FEXServer told us not to write any data
// other values: Code map writing is active
std::optional<int> CodeMapFD;
CodeMapOpener& FileOpener;
};
class AbstractCodeCache {
public:
virtual ~AbstractCodeCache() = default;
/**
* Computes a unique identifier for the referenced binary file to be used for
* generating the code map.
* This identifier is independent of FEX build/runtime configuration and
* stable across FEX updates.
*/
virtual uint64_t ComputeCodeMapId(std::string_view Filename, int FD) = 0;
/**
* Loads a code cache from mapped memory and appends it to the current Core state.
* TODO: Optionally recompiles all contained code blocks at runtime for validation.
+5 -5
View File
@@ -42,9 +42,6 @@ enum OperatingMode {
using CodeRangeInvalidationFn = std::function<void(uint64_t start, uint64_t Length)>;
// Nested vector of guest block entrypoints
using InvalidatedEntryAccumulator = fextl::vector<fextl::vector<uint64_t>>;
using CustomIREntrypointHandler = std::function<void(uintptr_t Entrypoint, IR::IREmitter*)>;
using ExitHandler = std::function<void(Core::InternalThreadState* Thread)>;
@@ -139,10 +136,13 @@ public:
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) = 0;
virtual AbstractCodeCache& GetCodeCache() = 0;
virtual void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter>) = 0;
virtual void FlushAndCloseCodeMap() = 0;
FEX_DEFAULT_VISIBILITY virtual void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(
FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void
InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual void
+6 -8
View File
@@ -104,6 +104,10 @@ struct CPUState {
uint64_t avx_high[16][2];
uint64_t gregs[16] {};
uint64_t L1Pointer {};
uint64_t L1Mask {};
uint64_t callret_sp {};
uint64_t _pad1 {};
XMMRegs xmm {};
// Raw segment register indexes
@@ -116,8 +120,6 @@ struct CPUState {
uint64_t gs_cached {};
uint64_t fs_cached {};
uint8_t flags[48] {};
uint64_t callret_sp {};
uint64_t _pad1 {};
uint64_t mm[8][2] {};
// 32bit x86 state
@@ -247,6 +249,8 @@ static_assert(offsetof(CPUState, xmm) % 32 == 0, "xmm needs to be 256-bit aligne
static_assert(offsetof(CPUState, mm) % 16 == 0, "mm needs to be 128-bit aligned!");
static_assert(offsetof(CPUState, gregs[15]) <= 504, "gregs maximum offset must be <= 504 for ldp/stp to work");
static_assert(offsetof(CPUState, DeferredSignalRefCount) % 8 == 0, "Needs to be 8-byte aligned");
static_assert(offsetof(CPUState, L1Pointer) <= 504, "This needs to be <= 504 for ldp");
static_assert(offsetof(CPUState, L1Mask) == (offsetof(CPUState, L1Pointer) + 8), "These two variables are paired");
struct InternalThreadState;
@@ -349,13 +353,11 @@ struct JITPointers {
uint64_t ExitFunctionLinker {};
uint64_t ThreadStopHandlerSpillSRA {};
uint64_t ThreadPauseHandlerSpillSRA {};
uint64_t UnimplementedInstructionHandler {};
uint64_t GuestSignal_SIGILL {};
uint64_t GuestSignal_SIGTRAP {};
uint64_t GuestSignal_SIGSEGV {};
uint64_t SignalReturnHandler {};
uint64_t SignalReturnHandlerRT {};
uint64_t L1Pointer {};
uint64_t L2Pointer {};
/** @} */
@@ -368,8 +370,6 @@ struct JITPointers {
// Process specific
uint64_t LUDIV {};
uint64_t LDIV {};
uint64_t LUREM {};
uint64_t LREM {};
// Thread Specific
@@ -378,8 +378,6 @@ struct JITPointers {
* @{ */
uint64_t LUDIVHandler {};
uint64_t LDIVHandler {};
uint64_t LUREMHandler {};
uint64_t LREMHandler {};
/** @} */
} AArch64;
@@ -40,6 +40,7 @@ struct HostFeatures {
bool SupportsECV {};
bool SupportsWFXT {};
bool Supports3DNow {};
bool SupportsSSE4a {};
// Float exception behaviour
bool SupportsAFP {};
@@ -67,6 +67,10 @@ public:
return Config;
}
virtual uintptr_t GetThunkCallbackRET() const {
return 0;
}
protected:
SignalDelegatorConfig Config;
};
@@ -4,6 +4,7 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/vector.h>
@@ -80,7 +81,14 @@ private:
static_assert(!std::is_move_constructible_v<NonMovableUniquePtr<int>>);
static_assert(!std::is_move_assignable_v<NonMovableUniquePtr<int>>);
struct InternalThreadState : public FEXCore::Allocator::FEXAllocOperators {
// Store used for unaligned LDAXR*/STLXR* emulation.
struct UnalignedExclusiveStore {
uint64_t Addr;
uint64_t Store;
uint8_t Size;
};
struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCore::Allocator::FEXAllocOperators {
FEXCore::Core::CpuStateFrame* const CurrentFrame = &BaseFrameState;
FEXCore::Context::Context* const CTX;
@@ -101,6 +109,8 @@ struct InternalThreadState : public FEXCore::Allocator::FEXAllocOperators {
// This pointer is owned by the frontend.
FEXCore::SHMStats::ThreadStats* ThreadStats {};
UnalignedExclusiveStore ExclusiveStore;
///< Data pointer for exclusive use by the frontend
void* FrontendPtr;
@@ -109,6 +119,10 @@ struct InternalThreadState : public FEXCore::Allocator::FEXAllocOperators {
// The low address of the call-ret stack allocation (not including guard pages)
void* CallRetStackBase {};
uintptr_t JITGuardPage {};
uint64_t JITGuardOverflowArgument {};
FEXCore::UncheckedLongJump::JumpBuf RestartJump;
// BaseFrameState should always be at the end, directly before the interrupt fault page
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState {};
@@ -116,8 +130,9 @@ struct InternalThreadState : public FEXCore::Allocator::FEXAllocOperators {
alignas(FEXCore::Utils::FEX_PAGE_SIZE) uint8_t InterruptFaultPage[FEXCore::Utils::FEX_PAGE_SIZE];
};
static_assert(std::is_standard_layout_v<FEXCore::Core::InternalThreadState>);
static_assert(
(offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState)) < 4096,
"Fault page is outside of immediate range from CPU state");
static_assert((offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState)) <
FEXCore::Utils::FEX_PAGE_SIZE,
"Fault page is outside of immediate range from CPU state");
static_assert(sizeof(FEXCore::Core::InternalThreadState) == (FEXCore::Utils::FEX_PAGE_SIZE * 2));
} // namespace FEXCore::Core
@@ -84,6 +84,6 @@ private:
class SourcecodeResolver {
public:
virtual fextl::unique_ptr<SourcecodeMap> GenerateMap(const std::string_view& GuestBinaryFile, const std::string_view& GuestBinaryFileId) = 0;
virtual fextl::unique_ptr<SourcecodeMap> GenerateMap(std::string_view GuestBinaryFile, std::string_view GuestBinaryFileId) = 0;
};
} // namespace FEXCore::HLE
@@ -54,10 +54,6 @@ public:
virtual ~SyscallHandler() = default;
virtual uint64_t HandleSyscall(FEXCore::Core::CpuStateFrame* Frame, FEXCore::HLE::SyscallArguments* Args) = 0;
virtual SyscallABI GetSyscallABI(uint64_t Syscall) = 0;
virtual FEXCore::IR::SyscallFlags GetSyscallFlags(uint64_t Syscall) const {
return FEXCore::IR::SyscallFlags::DEFAULT;
}
SyscallOSABI GetOSABI() const {
return OSABI;
+2 -29
View File
@@ -3,31 +3,12 @@
#include <FEXCore/Utils/EnumOperators.h>
#include <compare>
#include <cstdint>
#include <cstring>
namespace FEXCore::IR {
enum class SyscallFlags : uint8_t {
DEFAULT = 0,
// Syscalldoesn't care about CPUState being serialized up to the syscall instruction.
// Means dead code elimination can optimize through a syscall operation.
OPTIMIZETHROUGH = 1 << 0,
// Syscall only reads the passed in arguments. Doesn't read CPUState.
NOSYNCSTATEONENTRY = 1 << 1,
// Syscall doesn't return. Code generation after syscall return can be removed.
NORETURN = 1 << 2,
// Syscall doesn't have any side-effects, so if the result isn't used then it can be removed.
NOSIDEEFFECTS = 1 << 3,
// Syscall doesn't return a result.
// Means the resulting register shouldn't be written (Usually RAX).
// Usually used with !NOSYNCSTATEONENTRY, so the syscall can modify CPU state entirely.
// Then on return FEXCore picks up the new state.
NORETURNEDRESULT = 1 << 4,
};
FEX_DEF_NUM_OPS(SyscallFlags)
// This enum of named vector constants are linked to an array in CPUBackend.cpp.
// This is used with the IROp `LoadNamedVectorConstant` to load a vector constant
// that would otherwise be costly to materialize.
@@ -96,15 +77,7 @@ enum IndexNamedVectorConstant : uint8_t {
struct SHA256Sum final {
uint8_t data[32];
[[nodiscard]]
bool operator<(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) < 0;
}
[[nodiscard]]
bool operator==(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) == 0;
}
[[nodiscard]] auto operator<=>(const SHA256Sum&) const noexcept = default;
};
typedef void ThunkedFunction(void* ArgsRv);
@@ -82,12 +82,15 @@ inline bool VirtualProtect(void* Ptr, size_t Size, ProtectOptions options) {
return ::VirtualProtect(Ptr, Size, prot, nullptr) == 0;
}
inline void VirtualName(const char*, void*, size_t) {}
#else
using MMAP_Hook = void* (*)(void*, size_t, int, int, int, off_t);
using MUNMAP_Hook = int (*)(void*, size_t);
FEX_DEFAULT_VISIBILITY extern MMAP_Hook mmap;
FEX_DEFAULT_VISIBILITY extern MUNMAP_Hook munmap;
FEX_DEFAULT_VISIBILITY extern void VirtualName(const char* Name, void* Ptr, size_t Size);
// All commit parameters are ignored here, they are unnecessary as Linux supports overcommit
@@ -12,8 +12,6 @@ struct InternalThreadState;
namespace FEXCore::ArchHelpers::Arm64 {
enum class UnalignedHandlerType {
///< Don't backpatch code, instead handle inside SIGBUS handler.
Paranoid,
///< Backpatch unaligned access to half-barrier based atomic.
HalfBarrier,
///< Backpatch unaligned access to non-atomic.
@@ -26,7 +24,7 @@ enum class UnalignedHandlerType {
* This is an OS agnostic handler where the frontend must provide FEXCore with the information necessary to know if this is safe.
* This does not check if the PC is within a JIT code buffer, the frontend must provide that safety with `CPUBackend::IsAddressInCodeBuffer`.
*
* @param ParanoidTSO If the unaligned fault needs to handled directly or can be backpatched.
* @param HandleType Type of TSO handling to use.
* @param ProgramCounter The location in memory for the instruction that did the access
* @param GPRs The array of GPRs from the signal context. This will be modified and the host context needs to be updated on signal return.
*
@@ -34,6 +32,6 @@ enum class UnalignedHandlerType {
* by. FEXCore will return a positive or negative offset depending on internal handling.
*/
[[nodiscard]]
FEX_DEFAULT_VISIBILITY std::optional<int32_t>
HandleUnalignedAccess(FEXCore::Core::InternalThreadState* Thread, UnalignedHandlerType HandleType, uintptr_t ProgramCounter, uint64_t* GPRs);
FEX_DEFAULT_VISIBILITY std::optional<int32_t> HandleUnalignedAccess(
FEXCore::Core::InternalThreadState* Thread, UnalignedHandlerType HandleType, uintptr_t ProgramCounter, uint64_t* GPRs, bool IsJIT = true);
} // namespace FEXCore::ArchHelpers::Arm64
@@ -54,6 +54,13 @@ namespace FEXCore {
return static_cast<T>(key) == 0; \
}
// Macro that defines a fmt formatter for a reasonable case where an enum
// is formatted as a purely integral type based on its underlying type.
#define FEX_DEFINE_ENUM_FMT_PASSTHROUGH(type) \
constexpr auto format_as(type t) { \
return FEXCore::ToUnderlying(t); \
}
// Equivalent to C++23's std::to_underlying.
template<typename Enum>
[[nodiscard]]
+7 -4
View File
@@ -4,8 +4,10 @@
#include <cstdint>
// It's longjump without glibc fortification checks.
namespace FEXCore::LongJump {
// Reimplementation of longjmp without glibc fortification checks.
// This is useful when false positives need to be avoided or when using
// a libc implementation that does not implement std::longjmp.
namespace FEXCore::UncheckedLongJump {
// JumpBuf definition needs to be public because the frontend needs to understand it.
#if defined(_M_ARM_64)
struct JumpBuf {
@@ -32,5 +34,6 @@ struct JumpBuf {
#endif
[[nodiscard]] FEX_DEFAULT_VISIBILITY uint64_t SetJump(JumpBuf& Buffer);
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(JumpBuf& Buffer, uint64_t Value);
} // namespace FEXCore::LongJump
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(const JumpBuf& Buffer, uint64_t Value);
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC);
} // namespace FEXCore::UncheckedLongJump
@@ -0,0 +1,54 @@
// SPDX-License-Identifier: MIT
#pragma once
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/mman.h>
#include <sys/user.h>
#include <sys/prctl.h>
#ifndef PR_SET_VMA
#define PR_SET_VMA 0x53564d41
#endif
#ifndef PR_SET_VMA_ANON_NAME
#define PR_SET_VMA_ANON_NAME 0
#endif
#ifndef PR_GET_MEM_MODEL
#define PR_GET_MEM_MODEL 0x6d4d444c
#endif
#ifndef PR_SET_MEM_MODEL
#define PR_SET_MEM_MODEL 0x4d4d444c
#endif
#ifndef PR_SET_MEM_MODEL_DEFAULT
#define PR_SET_MEM_MODEL_DEFAULT 0
#endif
#ifndef PR_SET_MEM_MODEL_TSO
#define PR_SET_MEM_MODEL_TSO 1
#endif
#ifndef PR_GET_COMPAT_INPUT
#define PR_GET_COMPAT_INPUT 0x63494e50
#endif
#ifndef PR_SET_COMPAT_INPUT
#define PR_SET_COMPAT_INPUT 0x43494e50
#endif
#ifndef PR_SET_COMPAT_INPUT_DISABLE
#define PR_SET_COMPAT_INPUT_DISABLE 0
#endif
#ifndef PR_SET_COMPAT_INPUT_ENABLE
#define PR_SET_COMPAT_INPUT_ENABLE 1
#endif
#ifndef PR_GET_SHADOW_STACK_STATUS
#define PR_GET_SHADOW_STACK_STATUS 74
#endif
#ifndef PR_LOCK_SHADOW_STACK_STATUS
#define PR_LOCK_SHADOW_STACK_STATUS 76
#endif
#ifndef PR_SHADOW_STACK_ENABLE
#define PR_SHADOW_STACK_ENABLE (1ULL << 0)
#endif
#endif // ifndef _WIN32
Loaded 100 of 727 files, more files were not shown because too many files have changed in this diff. Show more