Compare commits

...
Author SHA1 Message Date
Ryan Houdek ba1b4744c5 Docs: Update for release FEX-2512 2025-12-05 16:11:40 -08:00
Ryan Houdek bd13c02451 Merge pull request #5092 from Sonicadvance1/12
SyscallsSMCTracking: Workaround assert in ELF mapping
2025-12-05 15:44:09 -08:00
Ryan Houdek 5902b175f9 SyscallsSMCTracking: Workaround assert in ELF mapping
The ELF tracking thing has an expectation that only portions of ELF
files that are described in the program headers will be mapped
executable. This doesn't hold true as programs will remap random
portions of ELF files as executable. In the case that this occurs, don't
assert out and instead print a warning.

This was discovered as Node.js remaps a portion of itself executable
that isn't described as such in the program headers. I also have a local
unittest that exposes the same problem. I had discovered same problem in
some other program with #5038.

Also fixes a bug where sometimes completely anonymously mapped
executable sneak in and cause a crash, which is kind of silly.
2025-12-05 14:50:58 -08:00
Ryan Houdek d0e47f9073 Merge pull request #5096 from neobrain/feature_fexofflinecompiler
CodeCache: Implement offline compiler for cache generation
2025-12-05 14:47:58 -08:00
Ryan Houdek bd7215d36f Merge pull request #5093 from Sonicadvance1/13
SteamRT4: Adds support for building the steam depot
2025-12-05 14:39:30 -08:00
Ryan Houdek f3f134f9de FEXServer: Add support for FEX Logging control
Always enables FEXServer log thread when built for Steam so that clients
can be controlled with `STEAM_FEX_LOG=1`. FEX logs will then always go
to FEXServer and those can get directed to wherever pressure-vessel
chooses.
2025-12-04 14:06:18 -08:00
Ryan Houdek cdbf5d57bd Steam: Adds FEXServerManager
This is fairly simple. Needs to be installed alongside FEXServer, so
that it can start it in a portable config.

- Starts a FEXServer
- Tells pressure-vessel when FEXServer is ready
- Keeps FEXServer alive with the `watch_fd` as long as the process lives
- Listens for pressure-vessel to be shutting down
- Exits once pressure-vessel exits, also letting FEXServer shutdown if
  no FEX instances are alive.
2025-12-04 14:06:17 -08:00
Ryan Houdek 1ae50bd670 FEXServerClient: Split out Connect and Start
Allow an optional watch_fd to be passed to FEXServer.
2025-12-04 14:02:43 -08:00
Ryan Houdek d28c9f9843 FEXServerClient: Disallow abstract named sockets under Steam
We don't want clients connecting to random sockets.
2025-12-04 14:02:43 -08:00
Ryan Houdek fe7d52aa78 github/steamrt4: Add artifacts 2025-12-04 14:02:43 -08:00
Ryan Houdek fc0907f8c1 Steam: Add FEXCompatTool 2025-12-04 14:02:43 -08:00
Ryan Houdek e57678d7a6 Config: Move Steam configs into config system 2025-12-04 14:02:43 -08:00
Ryan Houdek 45e594e806 Utils/StringUtils: Add in-place token replace helper 2025-12-04 14:02:43 -08:00
Ryan Houdek 87e7a0effa CMake: Don't install test thunk if tests aren't enabled 2025-12-04 14:02:43 -08:00
Ryan Houdek 4fd1a35b2a CMake: Disable some installs when building for Steam 2025-12-04 14:02:42 -08:00
Ryan Houdek c460cf0678 Merge pull request #5097 from wcampbell-nv/cmdline-map
Remap /proc/pid/cmdline with PR_SET_MM_MAP
2025-12-04 14:02:19 -08:00
Tony Wasserka 983802da61 CodeCache: Add offline compiler for generating caches 2025-12-04 19:16:51 +01:00
Tony Wasserka 49273e0d59 CodeCache: Add workaround for clang-15's broken std::piecewise_construct 2025-12-04 19:16:51 +01:00
Tony Wasserka 76b8459cdc CodeCache: Zero-initialize FileId in ExecutableFileInfo 2025-12-04 19:16:51 +01:00
Tony Wasserka e269eb6f65 ELFCodeLoader: Add helper interfaces 2025-12-04 19:16:51 +01:00
Tony Wasserka 7b4774f375 ELFCodeLoader: Add option to skip interpreter loading 2025-12-04 19:16:51 +01:00
Tony Wasserka 70e9a25112 Syscalls: Move m(un)map to a dedicated interface 2025-12-04 19:16:51 +01:00
Tony Wasserka 9fb83ea56a Windows/CRT: Implement lseek 2025-12-04 19:16:51 +01:00
LC 8bb3398376 Merge pull request #5101 from Sonicadvance1/16
SVE256: Fixes AVX scalar round with insert
2025-12-04 11:24:49 -05:00
Will Campbell d42fbb3d4d Use LoadFileToBuffer + cleanup 2025-12-04 07:54:45 -08:00
Tony Wasserka 53b2245dc1 Merge pull request #5098 from Sonicadvance1/14
docs: Update ProgrammingConcerns
2025-12-04 09:52:59 +00:00
Ryan Houdek db14975828 InstcountCI: Update 2025-12-04 01:45:32 -08:00
Ryan Houdek a5139d2710 unittests/ASM: Adds test case for AVX scalar round with insert bug 2025-12-04 01:44:12 -08:00
Ryan Houdek 9f584c8014 SVE256: Fixes AVX scalar round with insert
We were using the incorrect source registers on SVE256 implementation of
these instructions.

Fixes #5100
2025-12-04 01:43:07 -08:00
Tony Wasserka 6f1b98fb52 docs: Clean up ProgrammingConcerns 2025-12-04 10:18:26 +01:00
Ryan Houdek c258a90505 docs: Update ProgrammingConcerns
Disallow all APIs that touch `FILE`, they all allocate memory that we
don't control.
2025-12-03 20:32:33 -08:00
Will Campbell eb47ef43a7 Read from a fd rather than a FILE 2025-12-03 19:59:13 -08:00
Will Campbell 1876d6b923 Remap cmdline 2025-12-03 18:15:33 -08:00
Ryan Houdek 90c8fcf393 Merge pull request #5094 from pmatos/fix/issue5084
Set current code block in x87 pass
2025-12-02 14:43:54 -08:00
Ryan Houdek 6af90575e9 Merge pull request #5095 from neobrain/feature_serialize_relocations
JIT: Add support for serializing relocations
2025-12-02 14:43:42 -08:00
Tony Wasserka 994613260c CodeCache: Support reverse application of relocations
This allows code to be serialized consistently across runs.
2025-12-02 22:27:53 +01:00
Tony Wasserka 1430fa8220 CodeCache: Move ApplyCodeRelocations 2025-12-02 18:38:59 +01:00
Tony Wasserka 33e06058c6 JIT: Move ApplyRelocations to CodeCache 2025-12-02 18:38:59 +01:00
Tony Wasserka b032d1e1f7 JIT: Make relocations relative to guest base before serialization
This ensures consistency of generated code caches across multiple runs.
2025-12-02 18:38:59 +01:00
Tony Wasserka d6b43b1fe6 Dispatcher: Add public interface to query ExitFunctionLinkerAddress 2025-12-02 17:56:56 +01:00
Paulo Matos fc771c8683 asm_tests: Set current code block in x87 pass 2025-12-02 15:33:08 +01:00
Paulo Matos f91ac09f87 Set current code block in x87 pass
This resets the constant pool in IREmit used by SelectAddressMode().

Fixes #5084.
2025-12-02 15:33:08 +01:00
Ryan Houdek e4fa399412 Merge pull request #5091 from neobrain/feature_better_relocations
JIT: Prepare FEX relocations for code caching
2025-12-01 13:21:29 -08:00
Tony Wasserka 952e949e10 JIT: Clean up block tail writing code 2025-12-01 20:13:51 +01:00
Tony Wasserka 3d093d66fb JIT: Change relocation offset base to CodeBuffer start 2025-12-01 20:13:51 +01:00
Tony Wasserka 05fe2893c7 JIT: Emit relocation from IROP_THUNK 2025-12-01 20:13:51 +01:00
Tony Wasserka 6fc17294b6 JIT: Add relocation for guest RIP stored in jump thunks 2025-12-01 20:13:51 +01:00
Tony Wasserka 6607921bee JIT: Add relocation for guest RIP stored in block tail 2025-12-01 20:13:51 +01:00
Tony Wasserka 0b0793438f JIT: Add relocation for constants relative to the guest entrypoint 2025-12-01 20:13:51 +01:00
LC 3dd591e760 Merge pull request #5088 from Sonicadvance1/11
Linux: Disable io_uring
2025-12-01 14:05:56 -05:00
LC f290d2f899 Merge pull request #5090 from neobrain/refactor_3waycomp
IR: Replace hand-written operators with three-way comparison
2025-12-01 14:05:04 -05:00
Tony Wasserka a676ad7193 JIT: Add explicit padding to relocation descriptors
This ensures zero-initialization, which is required to make code cache
generation produce consistent results.

Also consolidated header fields.
2025-12-01 19:28:39 +01:00
Tony Wasserka 096c408ef6 IR: Replace hand-written operators with three-way comparison 2025-12-01 17:10:06 +01:00
Ryan Houdek 00b65f76b6 Linux: Disable io_uring
This allows passing around `epoll_event` structs which can't be
rewritten due to queues being managed by userspace.
2025-11-30 15:53:34 -08:00
Ryan Houdek 39fb266282 Merge pull request #5087 from neobrain/refactor_simpler_calls
OpcodeDispatcher: Simplify convoluted logic for computing call offsets
2025-11-28 10:10:22 -08:00
Ryan Houdek 3969d0ac78 Merge pull request #5086 from neobrain/fix_invalid_iterators
LookupCache: Fix use of invalidated iterators
2025-11-28 09:55:20 -08:00
Tony Wasserka 439c6bb3c0 OpcodeDispatcher: Simplify convoluted logic for computing call offsets 2025-11-28 11:32:11 +01:00
Tony Wasserka 5cedbf9d34 LookupCache: Fix use of invalidated iterators 2025-11-28 11:29:46 +01:00
Tony Wasserka 427b235eb5 Merge pull request #5071 from Sonicadvance1/2
Github: Add a steamrt4 builder
2025-11-28 08:59:46 +00:00
LC 92d5ba580f Merge pull request #5085 from Sonicadvance1/10
HostFeatures: Extend LRCPC2 errata to more CPUs
2025-11-28 00:41:50 -05:00
Ryan Houdek bc2f331c8b HostFeatures: Extend LRCPC2 errata to C1 Ultra/Premium 2025-11-27 21:29:17 -08:00
Ryan Houdek ca58aef676 HostFeatures: Extend LRCPC2 errata to V3AE 2025-11-27 21:18:51 -08:00
LC 379dc405f6 Merge pull request #5080 from Sonicadvance1/9
FEXInterpreter: Fixes crash with code maps
2025-11-27 20:42:20 -05:00
LC d74b5c42da Merge pull request #5079 from Sonicadvance1/8
FEXServer: Add support for a `wait_fd`
2025-11-27 20:41:22 -05:00
LC 6004971439 Merge pull request #5077 from Sonicadvance1/6
Scripts/InstallFEX: Fixes two issues
2025-11-27 20:39:07 -05:00
LC a12b8927bc Merge pull request #5076 from Sonicadvance1/5
Utils/WritePriorityMutex: Support being forkable
2025-11-27 20:38:37 -05:00
Ryan Houdek 57e23b289a Merge pull request #5081 from esullivan-nvidia/main
HostFeatures: Disable SupportsTSOImm9 for some CPUs
2025-11-27 16:22:40 -08:00
Tony Wasserka 8214ffccf0 Merge pull request #5083 from bylaws/wheiwofjsd
JIT: Fix indirect delinker branch distance
2025-11-27 15:02:47 +00:00
Billy Laws 547135dc2d JIT: Fix indirect delinker branch distance
This is in insts not bytes.
2025-11-27 14:42:56 +00:00
Ryan Houdek 9c72113161 Merge pull request #5070 from wcampbell-nv/cmdline
Reflect application changes to argv[0] in /proc/self/cmdline
2025-11-26 19:50:22 -08:00
esullivan 8cc967fa22 HostFeatures: Disable SupportsTSOImm9 for some CPUs
This change avoids using the LDAPUR instruction with CPUs that are know to be
impacted by an ARM CPU errata that results in poor performance.
2025-11-26 21:42:29 -06:00
Ryan Houdek 4bd30bb72d HostFeatures: Fixes bug in HostFeatures where simulator doesn't support new things
We now have a machine in CI that requires this.
2025-11-26 15:43:37 -08:00
Ryan Houdek 5eeb4dabbd Github: Add a steamrt4 builder
This ensures we don't break downstream projects.
2025-11-26 14:18:08 -08:00
Tony Wasserka a27c4b3860 Merge pull request #5078 from Sonicadvance1/7
CPUID: Fixes regression from #5033
2025-11-26 15:19:51 +00:00
Ryan Houdek dfee08f74f FEXInterpreter: Fixes crash with code maps
When the realpath of a program path can't be resolved, we weren't setting
the config option. This was cascading to be a crashing in codemaps where
it was unconditionally using the optional value (with assert checks),
and causing things to crash.

Pass in the path that can't be resolved to work around a crash in PV
that can happen.
2025-11-25 16:48:47 -08:00
Will Campbell 0e2629bdd4 Address review feedback 2025-11-25 14:09:29 -08:00
Ryan Houdek 05c8630b07 FEXServer: Add support for a wait_fd
This was a requested feature. To make sure that FEXServer is running and
managed by a parent process, we need to have a way to tell FEXServer to
keep alive without any FEX clients. The best way to do this is to pass
FEXServer a Pipe (like FEX does when a client starts it), but instead of
FEXServer signaling to FEXInterpreter that it's ready. FEXServer listens
to the pipe to see if the management process is still alive.

The expectation here is that the management process passes FEXServer the
read end of a pipe, and when the management software is done (or gets
killed by the kernel!) then the write end of the pipe is closed, and
FEXServer naturally closes (As long as there's no FEX processes
remaining).
2025-11-25 12:52:16 -08:00
Ryan Houdek 1cffe618d2 CPUID: Fixes #5033
This leaf changed to being non-constant on that PR since CPUID function
1h returns APICID now.
2025-11-25 11:25:36 -08:00
Ryan Houdek 98c7bb23b5 FEX/InstallFEX: Fixes issue of installing without software-properties-common
Checks to see if the package is installed first before trying to use it.
Fixes an issue where fresh users don't have this package installed and
the script fails.
2025-11-24 14:41:48 -08:00
Ryan Houdek 922853cee1 Scripts/InstallFEX: Fixes #4972
Makes sure that stderr output doesn't cause weird interactions with
FEXRootFSFetcher.
2025-11-24 14:40:56 -08:00
Tony Wasserka a251e61859 Merge pull request #5075 from Sonicadvance1/4
Config: Document the new `FEX_APP_CACHE_LOCATION` option
2025-11-24 20:24:59 +00:00
Ryan Houdek 9c19799023 Config: Document the new FEX_APP_CACHE_LOCATION option
I forgot to document this in the man page.
2025-11-24 12:07:54 -08:00
Ryan Houdek f423b110a8 Utils/WritePriorityMutex: Support being forkable
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.

This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
2025-11-24 11:50:20 -08:00
Ryan Houdek 6fd471e652 Merge pull request #5073 from discapes/unsquashfs-deco-fix
Support detecting unsquashfs>4.7.0 decompressors
2025-11-24 09:05:08 -08:00
Tony Wasserka 3d69029d33 Merge pull request #5072 from Sonicadvance1/3
Minor fixes
2025-11-24 11:24:07 +00:00
Miika Tuominen 2258f2e424 Support detecting unsquashfs>4.7.0 decompressors 2025-11-22 14:38:59 +02:00
Ryan Houdek cca5a68e20 SHMStats: Add missing header 2025-11-21 18:01:37 -08:00
Ryan Houdek da5c9bff68 Async: Add missing header. 2025-11-21 18:01:33 -08:00
Ryan Houdek 5b87f0699b pidof: Switch to using ranges 2025-11-21 18:01:28 -08:00
Will Campbell a67fe561a1 Reflect application changes to argv[0] in /proc/self/cmdline 2025-11-21 16:00:57 -08:00
Tony Wasserka e2f4065376 Merge pull request #5065 from Sonicadvance1/1
FEXCore/CodeCache: Move spin-loop over to a WFE loop
2025-11-21 14:21:45 +01:00
Ryan Houdek d3bf87f4f4 Merge pull request #4985 from Sonicadvance1/fex-atomic
Support (downstream) kernel-side unaligned atomic handling (The rebase sequel)
2025-11-20 17:09:29 -08:00
Ryan Houdek 6d351ec47f FEXCore/CodeCache: Moves spin-loop in to a WFE loop
Saves power and responds faster. Pass in the atomic to `WaitPred` with
the predicate checking if the buffer has been flushed yet. Same
behaviour as previous code but more efficient on our hardware.
2025-11-20 14:23:01 -08:00
Ryan Houdek 4dc1dd2511 FEXCore/Utils/SpinWaitLock: Adds WaitPred for waiting on a predicate
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.

We will need a predicated version for the next commit.
2025-11-20 14:23:01 -08:00
Ryan Houdek 5205ae40fa Merge pull request #5062 from pmatos/fix/address-size-handle
Implement address size modifier handling in CMPSOp and SCASOp
2025-11-20 12:19:22 -08:00
Ryan Houdek 32f1dcde7e Merge pull request #4906 from neobrain/feature_code_maps
CodeCache: Introduce code maps
2025-11-20 12:11:16 -08:00
Ryan Houdek 6772581c53 Arm64ec: Print a log when kernel unaligned atomics are used
Not having this in FEXInterpreter as it is too spammy in the general
case.
2025-11-20 12:05:55 -08:00
Ryan Houdek 387201815b Config: Add option to enable or disable kernel backpatchin on unaligned atomic
Default to enabled because this is the config we expect by default.
In the future will get some benchmarking in various games like
Assassin's Creed, and Call of Duty.
2025-11-20 12:04:59 -08:00
Billy Laws 24d61e1125 Windows: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Billy Laws 04259f031d FEXLoader: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Tony Wasserka 6403da3715 LinuxSyscalls: Shield code map FD from guest access
This prevents chromium/CEF from closing the FD.
2025-11-20 19:13:18 +01:00
Tony Wasserka 90cb76312c LinuxSyscalls: Implement code map writing for future code caching 2025-11-20 19:13:18 +01:00
Tony Wasserka b6cff01abb FEXServer: Add support for querying code maps
Managing code maps in FEXServer rather than in FEXInterpreter makes it
easier to handle multiple concurrent processes sharing code caches for
the main executable and libraries.
2025-11-20 19:13:18 +01:00
Tony Wasserka b34b711161 CodeCache: Add interfaces to describe and generate code maps
Code maps describe per-binary metadata used to generate caches. Currently,
this includes compiled block offsets and loaded shared libraries.
2025-11-20 19:13:18 +01:00
Ryan Houdek 709d767d61 FEX/Config: Allow override of cache location
This will be used.
2025-11-20 18:56:22 +01:00
Tony Wasserka e075916154 Config: Add interface to query cache directory 2025-11-20 18:56:22 +01:00
Paulo Matos de10154f29 instcountci: Implement address size modifier handling in CMPSOp and SCASOp for 64bits 2025-11-20 13:42:19 +01:00
Paulo Matos ba71e79e54 asm_tests: Implement address size modifier handling in CMPSOp and SCASOp 2025-11-20 13:42:19 +01:00
Paulo Matos 2cc70b8051 Implement address size modifier handling in CMPSOp and SCASOp for 64bits
A few games were generating "Can't handle adddress size".
I implemented 0x67 prefix handling for CMPSOp and SCASOP and improved
the error messages for the remainder. This will implement the address
modifier on 64bit systems, and keep issuing an error on 32bits.
2025-11-20 13:42:19 +01:00
Paulo Matos 9d965f94de asm_tests: Add 32-bit CMPS/SCAS tests without address size override 2025-11-20 13:42:19 +01:00
Ryan Houdek 40c2db4744 Merge pull request #5006 from bylaws/fasterrrrr
Introduce two-pass code invalidation model
2025-11-19 17:41:01 -08:00
Billy Laws 9a7285dca4 Windows: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws 8c00ac78b1 Linux: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws cf4478eeee LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-11-20 00:38:03 +00:00
Billy Laws 85c8e7f1bb fextl: Wrap tsl::robin_set 2025-11-20 00:38:03 +00:00
Billy Laws efd95efb40 FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-11-20 00:38:03 +00:00
Billy Laws 99ad7ea45c LookupCache: Drop unused state frame argument for delinker cbs 2025-11-20 00:38:03 +00:00
Ryan Houdek aba0c57f73 Merge pull request #5067 from pmatos/fix/Nasm3
Fix movzx instruction syntax
2025-11-19 14:01:15 -08:00
Paulo Matos 8c4f6b648e Fix movzx instruction syntax
nasm 2.16 was happy with it but it generates a bunch of errors in nasm3.
The generated binaries remain the same.
2025-11-19 15:04:16 +01:00
Ryan Houdek 3b83bdd88d Merge pull request #5066 from neobrain/fix_async_asserts
Async: Adapt precondition checks when receiving FDs
2025-11-19 01:45:41 -08:00
Tony Wasserka 8d71e08b44 Async: Strengthen precondition check when receiving FDs
The sender might provide all requested message bytes but no FD. The receiver
interface has no simple way of indicating this scenario yet, so just assert
out for now to ensure it never happens in the first place.

If needed, this can be changed to return a new error code to indicate partial
read in the future.
2025-11-19 09:48:23 +01:00
Tony Wasserka c31063a8ef Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-19 09:35:52 +01:00
Tony Wasserka 28d101f1dd Revert "Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark"
This cherry-picked an unfinished patch that wasn't intended for merging.
2025-11-19 09:35:08 +01:00
Ryan Houdek aaef344ae3 Merge pull request #5039 from Sonicadvance1/warkwarkwarkwarkwarkwark
FEX: Moves FEX thunk callback function generation to the frontend
2025-11-18 14:40:53 -08:00
Ryan Houdek 42d0324304 FEX: Moves FEX thunk callback function generation to the frontend
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.

Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.

Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.
2025-11-18 14:15:04 -08:00
Ryan Houdek 2a0019347a Thunks: Adds FEX Thunk callback to VDSO
This isn't necessary a VDSO, but it is /always/ mapped in to every
process. Use it as such.
2025-11-18 14:08:43 -08:00
Ryan Houdek e0305ea1b9 Merge pull request #5056 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEX/VDSO: Fixes symbol lookup
2025-11-18 12:51:13 -08:00
Ryan Houdek 105ff47ae3 FEX/VDSO: Fixes symbol lookup
Fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to .dynsym where gcc sticks them in to .symtab. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

Peeled out of #5039
2025-11-18 12:19:51 -08:00
Ryan Houdek 5ee190a41e Merge pull request #5060 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Config: Expose GetConv members
2025-11-17 10:45:02 -08:00
Ryan Houdek cf37617c25 Merge pull request #5011 from pmatos/feat/opt-memcpyf80
Refactoring of storing code in x87 opt. stack pass
2025-11-17 10:25:39 -08:00
Tony Wasserka 2e9c8f0f51 FEXCore/Config: Expose GetConv members
From working branch commit 8244ca1666796267ce25741cdf1103eef4f7539d
`Make cache generation aware of FEX configuration`
2025-11-17 10:21:51 -08:00
Ryan Houdek 06c2319851 Merge pull request #5061 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Common: Adds the ability to override HostFeatures registers by config.
2025-11-17 10:18:26 -08:00
Ryan Houdek da0668c7cc Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
Async: Move file descriptor checks from read_some() to read()
2025-11-17 10:17:44 -08:00
Ryan Houdek b34df334cb Merge pull request #5053 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection
2025-11-17 10:16:59 -08:00
Tony Wasserka 9e9f2ccae1 Merge pull request #5050 from neobrain/refactor_new_config_getter
Config: Refactor value getter interface
2025-11-17 18:58:50 +01:00
Tony Wasserka 15b8f75730 Config: Drop unneeded namespaces from StringArrayType 2025-11-17 18:46:35 +01:00
Tony Wasserka 9d6b9aa574 Config: Allow reading config values in arbitrary C++ expressions
FEX_CONFIG_OPT can only be used as a standalone statement, which is
inconvenient for config values that are only used once. The new functions
(e.g. Get_DUMPIR()) can be used in conditions or other expressions.
2025-11-17 18:46:02 +01:00
Tony Wasserka 64724886af Config: Replace macro-based config readers with a C++ template 2025-11-17 18:46:02 +01:00
Paulo Matos c088369f4a instcountci: Refactoring of storing code in x87 opt. stack pass 2025-11-17 10:14:29 +01:00
Paulo Matos 39dbf46422 Refactoring of storing code in x87 opt. stack pass
Enables memcpy optimization of 80bit floats on reduced precision.

Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
2025-11-17 10:14:29 +01:00
Paulo Matos 3c1b0bb917 Add instcountci tests for 80bit memcpy for x87 instructions 2025-11-17 10:14:29 +01:00
Ryan Houdek 2227170dbb FEXCore/Common: Adds the ability to override HostFeatures registers by config.
This is going to be necessary for the offline compiler work.
Also allow FEXGetConfig to print the same registers in the correct
format for easy fetching.
2025-11-16 16:43:01 -08:00
Ryan Houdek e08f421e1c Convert Async assert to logman assert 2025-11-16 14:56:30 -08:00
Tony Wasserka 42c58c5420 Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-16 14:54:46 -08:00
LC 4afbdd9afb Merge pull request #5055 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fixes alloca use-after-free in `RecvMMsg`
2025-11-14 20:26:05 -05:00
Ryan Houdek 06b9e13904 LinuxSyscalls: Fixes alloca use-after-free in RecvMMsg
Easy enough fix, thanks to @OFFTKP for pointing this out.
2025-11-14 16:27:27 -08:00
Tony Wasserka 29473b43cb LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection 2025-11-14 15:23:42 -08:00
Ryan Houdek 0427d48b98 Merge pull request #5033 from Sonicadvance1/warkwark
CPUID: Fixes APICID for processor count calculation.
2025-11-14 11:31:26 -08:00
Tony Wasserka b228746f1d FEXInterpreter: Fix incorrect value assignment
Previously this was overriding the cached value in the local Value object.
The global SilentLog variable never got updated, so a stale value would be
used.
2025-11-14 11:32:11 +01:00
Tony Wasserka 0be8485116 Merge pull request #5047 from antonkesy/fix_formatting
Align code with clang-format
2025-11-14 09:25:38 +01:00
Tony Wasserka 5d0279ff08 LibraryForwarding/gen: Tiny cleanup 2025-11-14 09:14:45 +01:00
Ryan Houdek 5d908d902c Merge pull request #5045 from lioncash/long
JIT: Handle long ADR/ADRP
2025-11-13 11:10:10 -08:00
Ryan Houdek 3a014f80f2 Merge pull request #5044 from antonkesy/clean_up_scripts
Scripts: Clean-up
2025-11-13 11:09:56 -08:00
Ryan Houdek fbefd7855c Merge pull request #5048 from neobrain/fix_fexconfig_string_lists
FEXConfig: Fix string list handling
2025-11-13 11:07:50 -08:00
Tony Wasserka 11f9135be6 FEXConfig: Fix string list handling 2025-11-13 17:33:21 +01:00
Lioncache 7bb0ce810e JIT: Expand LongAddressGen() to handle movz+movk sequence
This is only ever used on the path where we'd want to handle something
like this (in EmitEntryPoint()), so we can just extend the long handler
type instead of introducing a new type to handle this.
2025-11-13 11:16:34 -05:00
Lioncache 993b832771 JIT: Handle long ADR/ADRP
Wires up the long address handler into the ADR/ADRP restart
handlers.
2025-11-13 09:40:39 -05:00
Anton Kesy 013ac1e627 Align code with clang-format
Automatically done by running:
`find . \( -path './External' -prune \) -o \
  \( -iname '*.cc' -o -iname '*.cpp' -o -iname '*.hpp' -o \
     -iname '*.h' -o -iname '*.c' \) -print | \
  xargs clang-format --style=file -i`
2025-11-13 13:29:49 +01:00
Anton Kesy d7977a02fa remove semicolon 2025-11-12 21:45:09 +01:00
LC 73a32ff22c Merge pull request #5043 from antonkesy/fix_typos
Docs: fix typo
2025-11-12 15:28:24 -05:00
Anton Kesy ee4ae5390b move function comment inside function 2025-11-12 21:22:26 +01:00
Anton Kesy f71db11035 remove excess whitespaces 2025-11-12 21:22:08 +01:00
Anton Kesy 9497288b97 remove unused imports 2025-11-12 21:21:55 +01:00
Anton Kesy a00260d801 remove unused variable 2025-11-12 21:21:28 +01:00
Anton Kesy cad48e07e4 fix comment indentation 2025-11-12 21:21:05 +01:00
Anton Kesy d91e8a4278 docs: fix typo 2025-11-12 21:07:18 +01:00
LC faf74eee90 Merge pull request #5037 from Sonicadvance1/warkwarkwarkwark
FEXCore: Fixes JITGuardPage calculation in a threaded environment
2025-11-12 15:00:00 -05:00
Ryan Houdek 0b52e1cd14 Merge pull request #5042 from neobrain/fix_base_inference
LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
2025-11-12 11:18:58 -08:00
Tony Wasserka b62890f136 LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
At runtime, glxtest is mapped as follows:
0x000055fd9a030000 0x000055fd9a034000 0x4000  0x0     r--p  glxtest
0x000055fd9a034000 0x000055fd9a038000 0x4000  0x3000  r-xp  glxtest
0x000055fd9a038000 0x000055fd9a039000 0x1000  0x6000  rw-p  glxtest
0x000055fd9a039000 0x000055fd9a03a000 0x1000  0x6000  rw-p  glxtest

The problem here is that the last two sections can't be distinguished solely
by their mmap parameters. This would cause the wrong base address to be
inferred for the last mapping. To fix this, we can be more permissive by
allowing multiple candidates to be returned.

In practice, this only affects non-code sections, so it's not a big issue
either way.
2025-11-12 19:54:45 +01:00
Tony Wasserka de1d37eef8 Merge pull request #5041 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwark
CodeEmitter: Removes a few spurious asserts
2025-11-12 14:34:16 +01:00
Ryan Houdek a57c557485 CodeEmitter: Removes a few spurious asserts
These are handled with restart.
2025-11-11 18:25:21 -08:00
LC 3c9f6c845b Merge pull request #5040 from Sonicadvance1/warkwarkwarkwarkwarkwarkwark
FEXCore: Remove usage of "remote atomic" xor
2025-11-11 21:20:18 -05:00
Ryan Houdek ff25e9a92e FEXCore: Remove usage of "remote atomic" xor
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
2025-11-11 17:16:03 -08:00
Ryan Houdek 6a60f72a9e FEXCore: Fixes JITGuardPage calculation in a threaded environment
While this worked great for the singular unit test. I remembered thatour
pool allocator returns the minimum working size asked for but will
return larger sizes if exact fitment couldn't occur.

Because we are dealing with guard pages, we need to return the full
buffer size to the "client" so they can tell the frontend where the
guard page actually lives. Otherwise the JIT will tell the frontend the
guard page is at the end of the requested size, blow past the limit,
and fault in a completely different location.

With a bit of logging I saw in a multithreaded environment that we were
basically always getting a larger requested buffer while Steam was
starting up.
2025-11-11 12:31:38 -08:00
Ryan Houdek 94b690df43 unittests/FEXLinuxTests: Adds cpu core count test to cpuid
Ensures cpuid core counts are reported correctly.
2025-11-11 11:20:26 -08:00
Ryan Houdek 94edbc3436 CPUID: Stop accidentally exposing the HTT bit
We don't support this.
2025-11-11 11:20:26 -08:00
Ryan Houdek 5eab1e559a CPUID: Fixes APICID for processor count calculation.
Primary fix here is returning the current CPU index in function 01h.
Intel Quartus uses this alongside affinity setting to check if all cores
can be used for its calculation. Since we had hardcoded apicid 0 here,
it assumed to only have one core and never generated worker threads.

Additional fix for apicid size. This is the size of the bitmask required
for apic ids, we weren't calculating this correctly at all. This mask is
a "maximum" number of APICs that the CPU reserves in power of two.
Say the core supports 256 APICs, but the processor only supports 16, or
any other combination.
2025-11-11 11:20:26 -08:00
Ryan Houdek 53db3ad6f2 Merge pull request #5036 from neobrain/fix_thunkgen_glibcxx_debug
CMake: Disable libstdc++'s debug mode when compiling thunkgen
2025-11-11 09:14:53 -08:00
Tony Wasserka 581f3263ed CMake: Disable libstdc++'s debug mode when compiling thunkgen
This allows the rest of the project to use _GLIBCXX_DEBUG.
2025-11-11 17:26:37 +01:00
LC 1e3c642be6 Merge pull request #5035 from pmatos/fix/gradual-mem-growth
Use gradual memory growth
2025-11-11 09:00:27 -05:00
LC 22c3cd553f Merge pull request #5034 from Sonicadvance1/warkwarkwark
FEXCore/Win32: Move WritePriorityMutex away from SRWLock
2025-11-11 08:59:29 -05:00
Paulo Matos e5743f8dae Use gradual memory growth
Use min instead of max, otherwise we are always using `MAX_STATS_SIZE`.
2025-11-11 09:23:40 +01:00
Ryan Houdek bddc2f227d FEXCore/Win32: Move WritePriorityMutex away from SRWLock
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.

Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.

While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.

Dark Souls Remastered before:
```
  $RDLck Time: 4.531100 ms/second (0.04 percent)
  $WRLck Time: 2.122560 ms/second (0.02 percent)
```

after:
```
  $RDLck Time: 1.441620 ms/second (0.01 percent)
  $WRLck Time: 0.963720 ms/second (0.01 percent)
```
2025-11-10 17:43:47 -08:00
Ryan Houdek 686c04ea93 Win32: IMplement Wake/Wait by address 2025-11-10 17:43:40 -08:00
Ryan Houdek b38369199e Merge pull request #4893 from Sonicadvance1/long_long_codebuffer_pages
FEX: Implements support for JIT CodeBuffer guard page restart
2025-11-10 13:48:18 -08:00
Ryan Houdek e862c904a9 FEX: Implements support for JIT CodeBuffer guard page restart
When the JIT CodeBuffer overflows, we will now catch accesses to the
guard page and longjump while restarting the JIT with a larger buffer
request.

Fixes #4877
2025-11-10 11:55:21 -08:00
Ryan Houdek 43d9384b1c FEXCore/JIT: Add a pool allocator that understands a guard page
The size asked for has its final page guarded. It's up to the code
asking for allocations to ensure it never uses the final page if
necessary.
2025-11-10 11:54:25 -08:00
Ryan Houdek cb9af0b86a SignalDelegator: Split out SIGSEGV handler
This needs to run before the TestCodeHarness's frontend handler.
2025-11-10 11:53:40 -08:00
Ryan Houdek b8c17a843c ArchHelpers: Adds helper to get pointers to PC and FPRs 2025-11-07 15:33:47 -08:00
Ryan Houdek 7ad7f181d7 FEXCore/LongJump: Add a way to manually load from a longjump
The frontends will need this when loading a longjump buffer in to a
context.
2025-11-07 15:33:47 -08:00
Ryan Houdek eb0bf55033 FEXCore/JIT: Move the JIT long jump buffer to internalthreadstate
This will be a TLS variable that needs to be read by the frontend.
2025-11-07 15:33:47 -08:00
Ryan Houdek f4e3e4ad30 Merge pull request #5031 from lioncash/catch
Externals: Update catch2 from 3.5.3 to 3.11.0
2025-11-07 10:40:49 -08:00
Lioncache 5ae82410cc Externals: Update catch2 from 3.5.3 to 3.11.0
Updates it to the most recent release.
2025-11-07 08:31:04 -05:00
Ryan Houdek 2febb524e9 Merge pull request #5028 from lioncash/fmtup
Externals: Update fmt to 12.1.0
2025-11-06 10:04:15 -08:00
Lioncache b9e452133c Externals: Update fmt to 12.1.0
Keeps fmt updated to its latest release.
2025-11-06 10:17:22 -05:00
LC 747ea0a1f7 Merge pull request #5027 from Sbte/pr/xxhash
Update xxhash to v0.8.3
2025-11-06 07:38:56 -05:00
Sven Baars f8c52ca34a Update xxhash to v0.8.3 2025-11-06 11:53:55 +01:00
Ryan Houdek 663fd5a98b Docs: Update for release FEX-2511 2025-11-05 14:12:52 -08:00
Ryan Houdek 93e58bc15e Merge pull request #5009 from pmatos/feat/stack-xchange-opt
f80 stack xchg optimization for fast path
2025-11-05 12:54:41 -08:00
Ryan Houdek 3ccdf6508e Merge pull request #5024 from neobrain/feature_jit_encoder_recovery
FEXCore/JIT: Add support for recovering from branch encoding failures
2025-11-05 09:27:27 -08:00
Tony Wasserka fd33cf1ce5 Merge pull request #5020 from Sonicadvance1/fix_allocator_bugs
FEXCore/Allocator: Fixes two bugs
2025-11-05 11:53:41 +01:00
Tony Wasserka 2bb64ad1c6 FEXCore/Allocator: Require caller to move unique_ptr into release workaround
This further isolates the workaround to the implementation by highlighting
at the call-site that ownership is moved away.
2025-11-05 11:29:38 +01:00
Paulo Matos e3de62058b instcountci: f80 stack xchg optimization for fast path 2025-11-05 11:20:02 +01:00
Paulo Matos 6ee9984280 f80 stack xchg optimization for fast path 2025-11-05 11:20:02 +01:00
Tony Wasserka e1df548ae9 FEXCore: Extend documentation on uses for UncheckedLongJump 2025-11-05 10:11:30 +01:00
Tony Wasserka 79a685c15e FEXCore: Rename LongJump to UncheckedLongJump
This better reflects the difference to std::longjmp.
2025-11-05 10:11:30 +01:00
Tony Wasserka a61ab2803c CodeEmitter: Drop noisy (un)likely attributes
General usage of these attributes is discouraged. Since this is not
instruction-level performance critical code, drop them.
2025-11-05 09:47:02 +01:00
Ryan Houdek 209ad27332 Code view 2025-11-05 09:47:02 +01:00
Ryan Houdek df08981475 unittests/ASM: Adds test for too large branch objects 2025-11-05 09:47:02 +01:00
Ryan Houdek 5ce6039a02 FEXCore/JIT: Supports restarting JIT in case of encoding failure
ARM64 branches have fairly small relative distances they can encode.
These can be +-1MB, or even +-32KB. The largest relative branch is
+-128MB, which we already set as an upper limit of our block JIT cache
size.

We have for a long time just compiled these without checking with the
expectation that things just happen to work. We didn't hit the asserts
so it was relatively low priority. Apparently now with Steam and a
MaxInst limit of 5000, we are now hitting an assert where we are
encoding too large of a range.

Implement support for long jumping from anywhere in the JIT for when a
long jump tries to be encoded and fails, allowing us to restart the JIT
at any moment. This is implemented as a long jump when this singular
feature could have gotten away with some sort of invasive check and
early exit path for two reasons. For one, that would be even more
invasive, effectively doing try-catch logic manually. And two, the next
step is supporting JIT buffer overflow for when our block size heuristic
fails.

This next step will mandate longjump on SIGSEGV (with cooperative
interaction with the frontend) from effectively /anywhere/ in the JIT.
One of the design goals of the CodeEmitter is that every code emission
function doesn't do a size remaining check to allow the compiler to do
some very effective optimization of emitting code blocks to memory (and
it works!).

But we lose the ability to sanely size check. When writing the emitter I
knew we were going to need to write this cooperative guard page handler,
and we're finally at a point where it needs to be done. This will be in
the next PR although.
2025-11-05 09:47:02 +01:00
Ryan Houdek 6c3fdf723a FEXCore/JIT: Ignore local encoding limit checks
These are guaranteed not to hit encoding distance limits, so we can
ignore the returns.
2025-11-05 09:47:02 +01:00
Ryan Houdek 1a617c1eb2 FEXCore/Dispatcher: Check encoding errors 2025-11-05 09:47:02 +01:00
Ryan Houdek 3a9b801400 FEXCore/VectorRegType: Trivial header fix 2025-11-05 09:47:02 +01:00
Ryan Houdek 6e663acdad Linux/BPFEmitter: Explicitly ignored encoding bool
We know these won't encode in errors.
2025-11-05 09:47:02 +01:00
Ryan Houdek ae1023bb7a unittests/Emitter: Explicitly ignore encoding bool
We know these won't encode in errors.
2025-11-05 09:47:02 +01:00
Ryan Houdek 9cf25e276d CodeEmitter: Return bool if Label instructions can't be encoded
Programming error if they aren't checked, as they will encode
incorrectly if they are too large for their respective instructions.
2025-11-05 09:47:02 +01:00
Ryan Houdek 8db3670ecc FEXCore: Moves longjump implementation from FEX frontend
This will be getting used by FEXCore in a bit.
2025-11-05 09:47:02 +01:00
Ryan Houdek baee367532 FEXCore/Allocator: Move memory leak to a unified location 2025-11-04 16:06:05 -08:00
Ryan Houdek 8da4e72d87 FEXCore/Allocator: Fixes bug where MAP_FIXED could overallocate
When MAP_FIXED is used, if it was larger than the VMA region it was
trying to fit in to, then it would overallocate, corruption memory
adjacent to the VMA region. This was due to a typo in the LiveRegion
range checking.

Fix the typo, add a unittest that tries to overallocate space. Would
assert out without this bug fix.
2025-11-04 16:06:05 -08:00
Ryan Houdek b4a84a2317 Allocator: Fixes false OOM issue in allocator
In the case that overlapping `MAP_FIXED` mmap functions were used, we
were incorrectly tracking the full mapped regions size as new
allocation. We instead need to track which pages have already been
previously allocated and only track those. Would behave like FEX was
running out of memory, but we were just mapping the same location many
times.

Adds a unittest to track this.
2025-11-04 16:06:04 -08:00
Ryan Houdek 43d6347212 FlexBitSet: Add TestAndSet helper 2025-11-04 16:06:04 -08:00
Ryan Houdek 438501e49c Merge pull request #4998 from Sonicadvance1/i_like_my_writes_quick_and_monitored
LookupCache: Convert mutex to new WritePriorityMutex
2025-11-04 16:02:44 -08:00
Ryan Houdek c034e99aaf LookupCache: Convert mutex to new WritePriorityMutex
Changes the single highly-contended lock in `FindBlock` to be a
read-lock.
2025-11-04 15:48:52 -08:00
Ryan Houdek d2d0ca2de9 FEXCore: Implement a write-priority mutex
Now that our Lookup cache mutex is no longer recursive, we can safely
use a shared_mutex instead. The problem with a c++ std::shared_mutex is
that it doesn't guarantee any form of priority, so tens of thousands of
read-locks per second can cause a writer to never acquire the lock, or
take too much time.

The bad news is that C++ doesn't provide us a primitive with
write-priority, so we need to construct our own that is still compatible
with Linux futex. So this is what we do.

- Windows: Uses an SRWLock instead.
  - Only way for WINE to provide us a futex fallback that priorities
    write-priority without stampeding.
2025-11-04 15:48:52 -08:00
LC cbe2b442b2 Merge pull request #5021 from Sonicadvance1/fix_dir_iter
pidof: Fixes another unexpected throw location
2025-11-04 15:10:07 -05:00
LC c5b1cd6e7d Merge pull request #5023 from Sonicadvance1/wark
Wow64: Disable AVX
2025-11-04 15:09:31 -05:00
Ryan Houdek 8d20d1dae3 Wow64: Disable AVX
It's unsupported.
2025-11-04 11:29:04 -08:00
Billy Laws 1f2d702c4e Merge pull request #5004 from pmatos/simp/removeAsFloat
Remove InterpretAsFloat from x87StackOptimizationPass
2025-11-04 10:09:37 +00:00
Ryan Houdek d2d35d0553 pidof: Fixes another unexpected throw location
Exit early if the directory goes away before the iterator is created.
2025-11-03 12:31:03 -08:00
LC 1e7f54dd7e Merge pull request #5019 from Sonicadvance1/fix_flexbitset
FEXCore/Allocator: Fixes FlexBitSet
2025-11-03 08:42:22 -05:00
Ryan Houdek 8430a2f7e6 FEXCore/Allocator: Fixes FlexBitSet
A couple things here, we were never returning the last searched element,
either the last or first depending on search direction.

Also the backward scan would return incorrect indexes in some cases.
Also scanning beyond its page bounds.

Additionally some minorly incorrect assertions.

Adds a new unit test that ensures that we can allocate in to every
location, and that we get the correct indexes back. Also allocated
within guarded pages to ensure it doesn't read outside the bounds.

Fixes a spurious crash in Ender Magnolia.
2025-11-02 18:37:00 -08:00
LC 326cde78e6 Merge pull request #5016 from Sonicadvance1/vulkan_and_gl_fight_tonight
Thunks: Fixes symbol conflict between GL and Vulkan
2025-11-02 13:10:50 -05:00
LC bbc2b0b42f Merge pull request #5017 from Sonicadvance1/describe_esr
ArchHelpers: Adds ESR name helper
2025-11-01 23:20:16 -04:00
Ryan Houdek 231a2c54aa ArchHelpers: Adds ESR name helper
Just helps when an unhandled ESR occurs, it was always a case of needing
to go in to the ARM ARM to decode it which was a bit of a pain. Add a
textual representation of it.
2025-11-01 18:48:24 -07:00
LC 36ae4cee73 Merge pull request #5015 from Sonicadvance1/assert_fix
LookupCache: Fixes assert
2025-11-01 19:52:33 -04:00
LC 62e5ee2201 Merge pull request #5014 from Sonicadvance1/remove_unused_ptrs
FEXCore/CoreState: Removes some unused pointers
2025-11-01 19:51:46 -04:00
Ryan Houdek b89ebd931e Thunks: Fixes symbol conflict between GL and Vulkan
Fixes crash that occurs in applications that use both GL and Vulkan,
Like UE5 Vulkan native games. Fixes Ender Magnolia.

The issue here is that UE5 loads libGL first, which initializes our
libGL thunks, setting its X11Manager's functions.

It then loads libvulkan, which calls our oninit constructor, which
because of the symbol conflict, calls in to the libGL thunk's host
functions to reinitialize its function pointers, never initializing the
Vulkan X11Manager's functions. It would then crash as soon as an X11
function was used.

Give them unique symbol names so we don't accidentally look up the
incorrect symbol.
2025-11-01 16:48:56 -07:00
Ryan Houdek d1b4ddaf61 InstcountCI: Update 2025-11-01 15:18:44 -07:00
Ryan Houdek 734a0b236b FEXCore/CoreState: Removes some unused pointers
NFC
2025-11-01 15:13:10 -07:00
Ryan Houdek 2229c04d4d LookupCache: Fixes assert
These two asserts could never fail, Add assert to the base allocation
instead.
2025-11-01 15:11:44 -07:00
Ryan Houdek 07cff27fa2 SpinWaitLock: Adds one-shot WFE helper 2025-11-01 13:58:22 -07:00
Ryan Houdek 2cf86998bc SpinWaitLock: Fix missing pragma 2025-11-01 13:58:22 -07:00
Ryan Houdek 6ed15a6fd6 Windows: Implement support for SRWLock shared
Exclusive was already implemented.
2025-11-01 13:58:21 -07:00
Ryan Houdek 7610243b0c Merge pull request #5010 from neobrain/fix_asahi_regression
Switch back to jemalloc to fix regression in muvm-based setups
2025-11-01 13:34:36 -07:00
LC e11349b577 Merge pull request #5013 from Sonicadvance1/drm_v6.17
IoctlEmulation: Update to v6.17
2025-11-01 16:17:20 -04:00
Ryan Houdek 9b425697cb IoctlEmulation: Update to v6.17
Nova isn't handled yet because the API is in flux, but it's in v6.17 so
track it.
2025-11-01 13:04:52 -07:00
Ryan Houdek 42596ff91e External/drm-headers: Update to v6.17 2025-11-01 13:04:05 -07:00
LC b3c2ff47f3 Merge pull request #5012 from Sonicadvance1/qemu_apple
FEX: Update CPUID and detect script for newer qemu
2025-11-01 15:56:53 -04:00
Ryan Houdek 75a0bc79be FEX: Update CPUID and detect script for newer qemu
QEmu 10.2 is going to expose MIDR with Apple's vendor ID with variant 0.
That's the best they can do because they don't can't pin threads to
particular cores. So give a string for it, and detect it in the  fit
script.
2025-11-01 12:46:13 -07:00
Tony Wasserka 5a002ad08d Revert "Merge pull request #4969 from Sonicadvance1/rpmalloc"
This reverts commit e1a45a2720, reversing
changes made to bd7edd8651.

The change rendered pressure-vessel non-functional on muvm-based setups
like Fedora Asahi Remix.
2025-10-30 15:22:17 +01:00
Tony Wasserka fbac6f86d1 Merge pull request #5008 from neobrain/fix_removed_option
FEXConfig: Fix crash caused by no longer recognized option
2025-10-29 21:10:46 +01:00
Tony Wasserka ca18bf2a3d FEXConfig: Fix crash caused by no longer recognized option 2025-10-29 20:45:38 +01:00
Ryan Houdek 199effdff7 Merge pull request #5007 from bylaws/fasterefdfdgf
JIT: Restore behaviour of emitting interrupt checks at every block entry
2025-10-28 17:46:18 -07:00
Billy Laws 8212f4b7fb JIT: Restore behaviour of emitting interrupt checks at every block entry
This is needed to handle suspend in infinite loops that occur as a
result of block-size constraints or indirect jumps. Fixes grow home.
2025-10-29 00:34:56 +00:00
Ryan Houdek e1a45a2720 Merge pull request #4969 from Sonicadvance1/rpmalloc
Switch over to rpmalloc instead of jemalloc.
2025-10-28 17:25:25 -07:00
Ryan Houdek bd7edd8651 Merge pull request #5005 from bylaws/oodsakj
Profiler: Fix missing include
2025-10-28 17:25:02 -07:00
Paulo Matos 70b6bc2bae Remove InterpretAsFloat from x87StackOptimizationPass
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.

If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
2025-10-28 11:26:28 +01:00
Ryan Houdek 75391bf834 External: Remove jemalloc (jemalloc_glibc still exists) 2025-10-27 12:05:08 -07:00
Ryan Houdek 63304a1d88 Windows: rpmalloc 2025-10-27 12:05:08 -07:00
Ryan Houdek 985bdf2b6c Switch over to rpmalloc instead of jemalloc.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.

In Bayonetta's title screen it went from 963MB down to 834MB resident.
2025-10-27 11:23:13 -07:00
Ryan Houdek b57ea83aea External: Add rpmalloc 2025-10-27 11:23:13 -07:00
225 changed files with 8635 additions and 4233 deletions

No files matched your search

+79
View File
@@ -0,0 +1,79 @@
name: steamrt4 build
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
jobs:
steamrt4_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, distrobox]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name : submodule checkout
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build
cmake -E make_directory ${{runner.workspace}}/build
# Setup everything required.
- name : distrobox setup
run: |
distrobox create -Y -i registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306 steamrt4 || true
distrobox upgrade steamrt4
distrobox enter --name steamrt4 -- sudo apt-get install -y \
git cmake ninja-build ccache \
lld clang \
libclang-dev llvm-dev \
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross \
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- name: Create Build Environment
run: distrobox enter --name steamrt4 -- cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
shell: bash
working-directory: ${{runner.workspace}}/build
run: distrobox enter --name steamrt4 -- cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE
- name: install
working-directory: ${{runner.workspace}}/build
shell: bash
env:
DESTDIR: ${{runner.workspace}}/install
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE -t install
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
overwrite: true
name: steamrt4_steampipe_depot
path: ${{runner.workspace}}/install/*
retention-days: 1
compression-level: 9
+14 -2
View File
@@ -33,6 +33,7 @@ option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling ca
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use for the FEXCore profiler (gpuvis, tracy)")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
option(BUILD_STEAM_SUPPORT "Builds FEX for integration into Steam" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
@@ -64,6 +65,10 @@ if (NOT MINGW_BUILD)
endif()
endif()
if (BUILD_STEAM_SUPPORT)
add_definitions(-DFEX_STEAM_SUPPORT=1)
endif()
if (ENABLE_FEXCORE_PROFILER)
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
@@ -476,13 +481,16 @@ add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
if (_M_ARM_64 AND NOT MINGW_BUILD)
if (_M_ARM_64 AND NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
# Binfmt_misc files must be installed prior to Source/ installs
add_subdirectory(Data/binfmts/)
endif()
add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
if (NOT BUILD_STEAM_SUPPORT)
add_subdirectory(Data/AppConfig/)
endif()
# Install the ThunksDB file
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS ${CMAKE_CURRENT_SOURCE_DIR}/Data/*.json)
@@ -582,6 +590,10 @@ if (BUILD_THUNKS)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
if (BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
set(FEX_VERSION_MAJOR "0")
set(FEX_VERSION_MINOR "0")
set(FEX_VERSION_PATCH "0")
+63 -37
View File
@@ -36,24 +36,31 @@ public:
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adr(ARMEmitter::Register rd, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded adr(ARMEmitter::Register rd, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
if (IsADRRange(Imm)) {
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
}
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void adr(ARMEmitter::Register rd, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded adr(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::ADR});
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void adr(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded adr(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
adr(rd, &Label->Backward);
return adr(rd, &Label->Backward);
} else {
adr(rd, &Label->Forward);
return adr(rd, &Label->Forward);
}
}
@@ -62,38 +69,53 @@ public:
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adrp(ARMEmitter::Register rd, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded adrp(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) {
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
}
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void adrp(ARMEmitter::Register rd, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded adrp(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::ADRP});
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void adrp(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded adrp(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
adrp(rd, &Label->Backward);
return adrp(rd, &Label->Backward);
} else {
adrp(rd, &Label->Forward);
return adrp(rd, &Label->Forward);
}
}
void LongAddressGen(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>());
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, const BackwardLabel* Label) {
const auto SLocation = reinterpret_cast<int64_t>(Label->Location);
const auto ULocation = std::bit_cast<uint64_t>(SLocation);
const int64_t Imm = SLocation - (GetCursorAddress<int64_t>());
const auto UImm = std::bit_cast<uint64_t>(Imm);
if (IsADRRange(Imm)) {
// If the range is in ADR range then we can just use ADR.
adr(rd, Label);
} else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
return adr(rd, Label);
}
if (IsADRPRange(Imm)) {
const int64_t ADRPImm = (SLocation & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
// If the range is in the ADRP range then we can use ADRP.
bool NeedsOffset = !IsADRPAligned(reinterpret_cast<uint64_t>(Label->Location));
uint64_t AlignedOffset = reinterpret_cast<uint64_t>(Label->Location) & 0xFFFULL;
const bool NeedsOffset = !IsADRPAligned(ULocation);
const uint64_t AlignedOffset = ULocation & 0xFFFULL;
// First emit ADRP
adrp(rd, ADRPImm >> 12);
@@ -102,23 +124,33 @@ public:
// Now even an add
add(ARMEmitter::Size::i64Bit, rd, rd, AlignedOffset);
}
} else {
LOGMAN_MSG_A_FMT("Unscaled offset too large");
FEX_UNREACHABLE;
return BranchEncodeSucceeded::Success;
}
// Stinky path, we need to load the address as a sequence of movz+movk+movk
movz(ARMEmitter::Size::i64Bit, rd, (UImm >> 32) & 0xFFFF, 32);
movk(ARMEmitter::Size::i64Bit, rd, (UImm >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, rd, UImm & 0xFFFF);
return BranchEncodeSucceeded::Success;
}
void LongAddressGen(ARMEmitter::Register rd, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::LONG_ADDRESS_GEN});
// Emit a register index and a nop. These will be backpatched.
// Emit a register index and two nops. These will be backpatched.
dc32(rd.Idx());
nop();
nop();
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void LongAddressGen(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
LongAddressGen(rd, &Label->Backward);
return LongAddressGen(rd, &Label->Backward);
} else {
LongAddressGen(rd, &Label->Forward);
return LongAddressGen(rd, &Label->Forward);
}
}
@@ -862,12 +894,6 @@ public:
}
private:
static constexpr Condition InvertCondition(Condition cond) {
// These behave as always, so it makes no sense to allow inverting these.
LOGMAN_THROW_A_FMT(cond != Condition::CC_AL && cond != Condition::CC_NV, "Cannot invert CC_AL or CC_NV");
return static_cast<Condition>(FEXCore::ToUnderlying(cond) ^ 1);
}
void and_(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint32_t n, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b001'0010'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, n, immr, imms);
+123 -64
View File
@@ -20,23 +20,31 @@ public:
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm);
}
void b(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded b(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void b(ARMEmitter::Condition Cond, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded b(ARMEmitter::Condition Cond, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void b(ARMEmitter::Condition Cond, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded b(ARMEmitter::Condition Cond, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
b(Cond, &Label->Backward);
return b(Cond, &Label->Backward);
} else {
b(Cond, &Label->Forward);
return b(Cond, &Label->Forward);
}
}
@@ -45,24 +53,32 @@ public:
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm);
}
void bc(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded bc(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void bc(ARMEmitter::Condition Cond, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded bc(ARMEmitter::Condition Cond, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void bc(ARMEmitter::Condition Cond, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded bc(ARMEmitter::Condition Cond, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
bc(Cond, &Label->Backward);
return bc(Cond, &Label->Backward);
} else {
bc(Cond, &Label->Forward);
return bc(Cond, &Label->Forward);
}
}
@@ -98,25 +114,32 @@ public:
UnconditionalBranch(Op, Imm);
}
void b(const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded b(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0001'01 << 26;
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
UnconditionalBranch(Op, Imm >> 2);
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void b(ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded b(ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::B});
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void b(BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded b(BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
b(&Label->Backward);
return b(&Label->Backward);
} else {
b(&Label->Forward);
return b(&Label->Forward);
}
}
@@ -126,25 +149,33 @@ public:
UnconditionalBranch(Op, Imm);
}
void bl(const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded bl(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b1001'01 << 26;
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
UnconditionalBranch(Op, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void bl(ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded bl(ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::B});
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void bl(BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded bl(BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
bl(&Label->Backward);
return bl(&Label->Backward);
} else {
bl(&Label->Forward);
return bl(&Label->Forward);
}
}
@@ -155,28 +186,35 @@ public:
CompareAndBranch(Op, s, rt, Imm);
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded cbz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0100 << 24;
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
CompareAndBranch(Op, s, rt, Imm >> 2);
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded cbz(ARMEmitter::Size s, ARMEmitter::Register rt, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
cbz(s, rt, &Label->Backward);
return cbz(s, rt, &Label->Backward);
} else {
cbz(s, rt, &Label->Forward);
return cbz(s, rt, &Label->Forward);
}
}
@@ -186,28 +224,35 @@ public:
CompareAndBranch(Op, s, rt, Imm);
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0101 << 24;
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
CompareAndBranch(Op, s, rt, Imm >> 2);
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::BC});
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
cbnz(s, rt, &Label->Backward);
return cbnz(s, rt, &Label->Backward);
} else {
cbnz(s, rt, &Label->Forward);
return cbnz(s, rt, &Label->Forward);
}
}
@@ -217,28 +262,35 @@ public:
TestAndBranch(Op, rt, Bit, Imm);
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded tbz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0110 << 24;
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
TestAndBranch(Op, rt, Bit, Imm >> 2);
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded tbz(ARMEmitter::Register rt, uint32_t Bit, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::TEST_BRANCH});
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void tbz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded tbz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
tbz(rt, Bit, &Label->Backward);
return tbz(rt, Bit, &Label->Backward);
} else {
tbz(rt, Bit, &Label->Forward);
return tbz(rt, Bit, &Label->Forward);
}
}
@@ -247,27 +299,34 @@ public:
TestAndBranch(Op, rt, Bit, Imm);
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded tbnz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0111 << 24;
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
}
TestAndBranch(Op, rt, Bit, Imm >> 2);
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, ForwardLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded tbnz(ARMEmitter::Register rt, uint32_t Bit, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::TEST_BRANCH});
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, 0);
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
}
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel* Label) {
[[nodiscard]] BranchEncodeSucceeded tbnz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel* Label) {
if (Label->Backward.Location) {
tbnz(rt, Bit, &Label->Backward);
return tbnz(rt, Bit, &Label->Backward);
} else {
tbnz(rt, Bit, &Label->Forward);
return tbnz(rt, Bit, &Label->Forward);
}
}
+78 -31
View File
@@ -586,6 +586,15 @@ concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegi
template<typename T>
concept IsQOrDRegister = std::is_same_v<T, QRegister> || std::is_same_v<T, DRegister>;
template<typename T>
concept IsLabel = std::is_same_v<T, ARMEmitter::ForwardLabel> || std::is_same_v<T, ARMEmitter::BackwardLabel> ||
std::is_same_v<T, ARMEmitter::BiDirectionalLabel> || std::is_same_v<T, ARMEmitter::ForwardLabel::Reference>;
enum class BranchEncodeSucceeded {
Success,
Failure,
};
// Whether or not a given set of vector registers are sequential
// in increasing order as far as the register file is concerned (modulo its size)
//
@@ -638,19 +647,25 @@ public:
// Bind a backward label to an address.
// Address that is bound is the current emitter location.
void Bind(BackwardLabel* Label) {
[[nodiscard]] bool Bind(BackwardLabel* Label) {
LOGMAN_THROW_A_FMT(Label->Location == nullptr, "Trying to bind a label twice");
Label->Location = GetCursorAddress<uint8_t*>();
// Always binds because it is only storing a location.
return true;
}
void Bind(const ForwardLabel::Reference* Label) {
[[nodiscard]] bool Bind(const ForwardLabel::Reference* Label) {
uint8_t* CurrentAddress = GetCursorAddress<uint8_t*>();
// Patch up the instructions
switch (Label->Type) {
case ForwardLabel::InstType::ADR: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
if (!IsADRRange(Imm)) {
// Can't bind.
return false;
}
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
@@ -662,7 +677,12 @@ public:
case ForwardLabel::InstType::ADRP: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
if (!(IsADRPRange(Imm) && IsADRPAligned(Imm))) {
// Can't bind.
return false;
}
Imm >>= 12;
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
@@ -672,11 +692,13 @@ public:
*Instruction = Inst;
break;
}
case ForwardLabel::InstType::B: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (!(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0))) {
// Can't bind.
return false;
}
Imm >>= 2;
uint32_t InstMask = 0x3FF'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
@@ -686,11 +708,13 @@ public:
break;
}
case ForwardLabel::InstType::TEST_BRANCH: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (!(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0))) {
// Can't bind.
return false;
}
Imm >>= 2;
uint32_t InstMask = 0x3FFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
@@ -704,7 +728,10 @@ public:
case ForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (!(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0))) {
// Can't bind.
return false;
}
Imm >>= 2;
uint32_t InstMask = 0x7'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
@@ -714,38 +741,44 @@ public:
break;
}
case ForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
const auto* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
const auto ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
const auto ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
const auto ImmInstThree = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[2]);
const auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
const auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
if (IsADRRange(ImmInstThree)) {
// If within ADR range from the third instruction, then we can emit NOP+NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
} else if (IsADRPRange(ImmInstOne)) {
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstThree) & 0x7FFF);
} else if (IsADRPRange(ImmInstTwo)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
// We can emit nop + nop + adrp
nop();
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstThree >> 12) & 0x7FFF);
} else {
// Not aligned, need nop + adrp + add
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
} else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstTwo & 0xFFF);
}
} else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
// Stinky path, we need to emit a movz+movk+movk sequence.
movz(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 32) & 0x7FFF, 32);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne) & 0xFFFF);
}
SetCursorOffset(OriginalOffset);
@@ -753,27 +786,41 @@ public:
}
default: LOGMAN_MSG_A_FMT("Unexpected inst type in label fixup");
}
return true;
}
// Bind a forward label to a location.
// This walks all the instructions in the label's vector.
// Then backpatching all instructions that have used the label.
void Bind(ForwardLabel* Label) {
[[nodiscard]] bool Bind(ForwardLabel* Label) {
bool Bound = true;
if (Label->FirstInst.Location) {
Bind(&Label->FirstInst);
Bound &= Bind(&Label->FirstInst);
}
for (auto& Inst : Label->Insts) {
Bind(&Inst);
Bound &= Bind(&Inst);
}
return Bound;
}
// Bind a bidirectional location to a location.
// Binds both forwards and backwards depending on how the label was used.
void Bind(BiDirectionalLabel* Label) {
[[nodiscard]] bool Bind(BiDirectionalLabel* Label) {
bool Bound = true;
if (!Label->Backward.Location) {
Bind(&Label->Backward);
Bound &= Bind(&Label->Backward);
}
Bind(&Label->Forward);
Bound &= Bind(&Label->Forward);
return Bound;
}
static constexpr Condition InvertCondition(Condition cond) {
// These behave as always, so it makes no sense to allow inverting these.
LOGMAN_THROW_A_FMT(cond != Condition::CC_AL && cond != Condition::CC_NV, "Cannot invert CC_AL or CC_NV");
return static_cast<Condition>(FEXCore::ToUnderlying(cond) ^ 1);
}
#include <CodeEmitter/VixlUtils.inl>
+1 -1
+1 -1
+1 -1
+5 -3
View File
@@ -74,9 +74,11 @@ add_compile_options($<$<COMPILE_LANGUAGE:CXX>:-fno-strict-aliasing> $<$<COMPILE_
add_subdirectory(Source/)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
if (NOT BUILD_STEAM_SUPPORT)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
endif()
if (BUILD_TESTING)
add_subdirectory(unittests/)
+9
View File
@@ -200,6 +200,15 @@ def print_man_environment_tail():
],
"''", True)
print_man_env_option(
"APP_CACHE_LOCATION",
[
"Allows the user to override where FEX stores and loads cache files",
"By default FEX will look in $XDG_CACHE_HOME/fex-emu/ or $HOME/.cache/fex-emu/",
"This will override the full path, trailing forward-slash is expected to exist",
],
"''", True)
def print_man_header():
header ='''.Dd {0}
.Dt FEX
+6 -4
View File
@@ -31,7 +31,6 @@ set (SRCS
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
@@ -66,6 +65,7 @@ set (SRCS
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/x87StackOptimizationPass.cpp
Utils/LongJump.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
@@ -201,8 +201,10 @@ add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_MAN_NAME}"
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
if (NOT BUILD_STEAM_SUPPORT)
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
endif()
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
@@ -289,7 +291,7 @@ AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
if (NOT MINGW_BUILD)
if (NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
install(TARGETS ${PROJECT_NAME}_shared
LIBRARY
DESTINATION ${CMAKE_INSTALL_LIBDIR}
+2
View File
@@ -4,6 +4,8 @@
#ifdef _M_X86_64
#include <xmmintrin.h>
#include <immintrin.h>
#else
#include <cstdint>
#endif
namespace FEXCore {
+19 -14
View File
@@ -30,14 +30,14 @@ class Context;
}
namespace FEXCore::Config {
namespace DefaultValues {
namespace detail {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) const P(type) P(enum) = P(default);
#define OPT_STR(group, enum, json, default) const std::string_view P(enum) = P(default);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#define OPT_STRENUM(group, enum, json, default) const uint64_t P(enum) = FEXCore::ToUnderlying(P(default));
#include <FEXCore/Config/ConfigValues.inl>
} // namespace DefaultValues
} // namespace detail
enum Paths {
PATH_DATA_DIR_LOCAL = 0,
@@ -134,7 +134,7 @@ public:
void Load();
template<typename T>
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<DefaultValues::Type::StringArrayType, T>)
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<StringArrayType, T>)
std::optional<T> GetConv(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
@@ -142,7 +142,7 @@ public:
}
const auto& Value = it->second;
LOGMAN_THROW_A_FMT(!std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(!std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
if (std::holds_alternative<T>(Value)) [[likely]] {
return std::get<T>(Value);
@@ -165,7 +165,7 @@ public:
private:
void MergeConfigMap(const LayerOptions& Options);
void MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value);
void MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value);
};
void MetaLayer::Load() {
@@ -181,7 +181,7 @@ void MetaLayer::Load() {
}
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value) {
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
@@ -193,7 +193,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](const DefaultValues::Type::StringArrayType& Value) {
const auto AddToMap = [&LookupMap](const StringArrayType& Value) {
for (const auto& EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == fextl::string::npos) {
@@ -209,7 +209,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
}
};
AddToMap(std::get<DefaultValues::Type::StringArrayType>(MetaEnvironment->second));
AddToMap(std::get<StringArrayType>(MetaEnvironment->second));
AddToMap(Value);
// Now with the two layers merged in the map
@@ -225,8 +225,8 @@ void MetaLayer::MergeConfigMap(const LayerOptions& Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto& it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV || it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<DefaultValues::Type::StringArrayType>(it.second));
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<StringArrayType>(it.second));
} else {
OptionMap.insert_or_assign(it.first, it.second);
}
@@ -423,7 +423,7 @@ bool Exists(ConfigOption Option) {
return Meta->OptionExists(Option);
}
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
std::optional<StringArrayType*> All(ConfigOption Option) {
return Meta->All(Option);
}
@@ -436,6 +436,12 @@ std::optional<T> GetConv(ConfigOption Option) {
return Meta->GetConv<T>(Option);
}
template std::optional<bool> GetConv(ConfigOption Option);
template std::optional<uint8_t> GetConv(ConfigOption Option);
template std::optional<int32_t> GetConv(ConfigOption Option);
template std::optional<uint32_t> GetConv(ConfigOption Option);
template std::optional<uint64_t> GetConv(ConfigOption Option);
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
@@ -491,13 +497,12 @@ template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t De
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List) {
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<DefaultValues::Type::StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option,
DefaultValues::Type::StringArrayType* List);
template void Value<StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
} // namespace FEXCore::Config
@@ -16,6 +16,13 @@
"Maximum number of instruction to store in a block"
]
},
"EnableCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable the code caching subsystem"
]
},
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
@@ -94,6 +101,13 @@
"Desc": [
"Scales the cycle counter on systems that have low frequencies."
]
},
"CPUFeatureRegisters": {
"Type": "str",
"Default": "",
"Desc": [
"Allows overriding cpu feature flags for manual testing"
]
}
},
"Emulation": {
@@ -429,6 +443,13 @@
"This is required to ensure a split-lock doesn't tear inside the process"
]
},
"KernelUnalignedAtomicBackpatching": {
"Type": "bool",
"Default": "true",
"Desc": [
"When the kernel unaligned atomic handler is enabled, use backpatching to reduce kernel context switches."
]
},
"VolatileMetadata": {
"Type": "bool",
"Default": "true",
+41 -9
View File
@@ -4,7 +4,6 @@
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -62,7 +61,7 @@ struct CustomIRResult {
, Data(Data) {}
};
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class CodeCache : public AbstractCodeCache {
@@ -73,12 +72,34 @@ public:
ContextImpl& CTX;
bool IsGeneratingCache = false;
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
void LoadData(Core::InternalThreadState&, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
void InitiateCacheGeneration() override {
IsGeneratingCache = true;
}
/**
* Applies a set of FEX relocations to the given code section.
*
* FEX relocations describe runtime-dependencies of FEX-generated code.
* When loading a code cache, they are used to move cached code to the
* dynamically chosen base address of the guest binary.
*
* Conversely, relocations are applied in reverse when writing code caches
* to ensure consistency across generation runs.
*
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
@@ -155,10 +176,20 @@ public:
return CodeCache;
}
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter> Writer) override {
CodeMapWriter = std::move(Writer);
}
void FlushAndCloseCodeMap() override {
if (CodeMapWriter) {
CodeMapWriter.reset();
}
}
void OnCodeBufferAllocated(const std::shared_ptr<CPU::CodeBuffer>&) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start,
uint64_t Length) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -225,14 +256,12 @@ public:
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CodeCache CodeCache;
fextl::unique_ptr<CodeMapWriter> CodeMapWriter;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl(const FEXCore::HostFeatures& Features);
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, const FEXCore::LookupCacheWriteLockToken& lk);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
// This is used as a replacement for the SMC writes in the mono callsite backpatcher that avoids atomic operations
@@ -269,7 +298,7 @@ public:
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator {"FEXMem_OpDispatcher"};
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator {"FEXMem_Frontend"};
FEXCore::Utils::PooledAllocatorVirtual CPUBackendAllocator {"FEXMem_CPUBackend"};
FEXCore::Utils::PooledAllocatorVirtualWithGuard CPUBackendAllocator {"FEXMem_CPUBackend"};
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const {
@@ -348,5 +377,8 @@ private:
bool MonoDetected = false;
std::atomic<uint64_t> MonoBackpatcherBlock;
std::mutex CodeBufferListLock;
fextl::vector<std::weak_ptr<CPU::CodeBuffer>> CodeBufferList;
};
} // namespace FEXCore::Context
@@ -105,9 +105,12 @@ constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public ARMEmitter::Emitter {
protected:
public:
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
protected:
FEXCore::Context::ContextImpl* EmitterCTX;
std::span<const ARMEmitter::Register> StaticRegisters {};
@@ -117,8 +120,6 @@ protected:
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
// Correlate an ARM register back to an x86 register index.
+1 -1
View File
@@ -400,7 +400,7 @@ namespace CPU {
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(*Buffer);
OnCodeBufferAllocated(Buffer);
return Buffer;
}
+2 -2
View File
@@ -81,7 +81,7 @@ namespace CPU {
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(CodeBuffer&) {};
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
@@ -161,7 +161,7 @@ namespace CPU {
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) = 0;
virtual void ClearCache() {}
+11 -8
View File
@@ -14,6 +14,7 @@ $end_info$
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Syscalls.h>
@@ -88,6 +89,7 @@ namespace ProductNames {
static const char ARM_Blizzard_M2Pro[] = "Apple Blizzard (M2 Pro)";
static const char ARM_Avalanche_M2Max[] = "Apple Avalanche (M2 Max)";
static const char ARM_Blizzard_M2Max[] = "Apple Blizzard (M2 Max)";
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_Ampere_1[] = "AmpereOne";
@@ -188,6 +190,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x61, 0x029, 1, ProductNames::ARM_Firestorm_M1Max}, // Apple Firestorm (M1 Max)
{0x61, 0x025, 1, ProductNames::ARM_Firestorm_M1Pro}, // Apple Firestorm (M1 Pro)
{0x61, 0x023, 1, ProductNames::ARM_Firestorm_M1}, // Apple Firestorm (M1)
{0x61, 0, 1, ProductNames::ARM_AppleSilicon}, // QEmu Apple Silicon
{0x41, 0xd8c, 1, ProductNames::ARM_C1Ultra}, // C1-Ultra
{0x41, 0xd90, 1, ProductNames::ARM_C1Premium}, // C1-Premium
@@ -441,10 +444,10 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
Res.eax = FAMILY_IDENTIFIER;
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(0 << 24); // Local APIC ID
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(GetCPUID() << 24); // Local APIC ID
Res.ecx = (1 << 0) | // SSE3
(CTX->HostFeatures.SupportsPMULL_128Bit << 1) | // PCLMULQDQ
@@ -507,7 +510,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
(1 << 25) | // SSE
(1 << 26) | // SSE2
(0 << 27) | // Self Snoop
(1 << 28) | // Max APIC IDs reserved field is valid
(0 << 28) | // (HTT) Max APIC IDs reserved field is valid
(1 << 29) | // Thermal monitor
(0 << 30) | // Reserved
(0 << 31); // Pending break enable
@@ -1094,9 +1097,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) con
(CTX->HostFeatures.SupportsCLZERO << 0); // CLZERO support
uint32_t CoreCount = Cores - 1;
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
((uint32_t)std::log2(CoreCount + 1) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
(std::bit_ceil(Cores) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
return Res;
}
+1 -1
View File
@@ -277,7 +277,7 @@ private:
// 0: Highest function parameter and ID
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 1: Processor info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 2: Cache and TLB info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 3: Serial Number(previously), now reserved
+347 -2
View File
@@ -1,12 +1,213 @@
// SPDX-License-Identifier: MIT
#include <Interface/Context/Context.h>
#include "Utils/SpinWaitLock.h"
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
#include <Interface/Core/JIT/Relocations.h>
#include <Interface/Core/LookupCache.h>
#include <FEXCore/Core/Thunks.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <git_version.h>
#include <xxhash.h>
#include <fstream>
namespace FEXCore {
#if __clang_major__ < 16
ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map, uint64_t FileId, fextl::string Filename)
: SourcecodeMap(std::move(Map))
, FileId(FileId)
, Filename(Filename) {}
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
std::string_view base_filename = FHU::Filesystem::GetFilename(std::string_view {MainExecutable.Filename});
if (FileId != 0xffff'ffff'ffff'ffff) {
return fextl::fmt::format("{}-{:016x}{}", base_filename, MainExecutable.FileId, AddNombSuffix ? "-nomb" : "");
}
return "";
}
fextl::map<CodeMapFileId, CodeMap::ParsedContents> CodeMap::ParseCodeMap(std::ifstream& File) {
fextl::map<CodeMapFileId, CodeMap::ParsedContents> Ret;
while (true) {
Entry Entry;
File.read(reinterpret_cast<char*>(&Entry), sizeof(Entry));
if (!File) {
break;
}
if (Entry.FileId == LoadExternalLibrary.FileId && Entry.BlockOffset == LoadExternalLibrary.BlockOffset) {
ExternalLibraryInfo Info;
File.read(reinterpret_cast<char*>(&Info), sizeof(Info));
fextl::string Filename;
std::getline(File, Filename, '\0');
// Align to 4-byte boundary
char Null[4];
File.read(Null, AlignUp(Filename.size() + 1, 4) - Filename.size() - 1);
if (!File) {
break;
}
Ret[Info.ExternalFileId].Filename = std::move(Filename);
} else if (Entry.FileId == SetExecutableFileId {}.Marker.FileId && Entry.BlockOffset == SetExecutableFileId {}.Marker.BlockOffset) {
CodeMapFileId ExecutableFileId;
File.read(reinterpret_cast<char*>(&ExecutableFileId), sizeof(ExecutableFileId));
if (!File) {
break;
}
Ret[ExecutableFileId].IsExecutable = true;
} else {
if (!Ret.contains(Entry.FileId)) {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
} else {
Ret[Entry.FileId].Blocks.insert(Entry.BlockOffset);
}
}
if (!File) {
break;
}
}
return Ret;
}
CodeMapWriter::CodeMapWriter(CodeMapOpener& Opener, bool OpenEagerly)
: Buffer(4096)
, FileOpener(Opener) {
if (OpenEagerly) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
}
CodeMapWriter::~CodeMapWriter() {
if (CodeMapFD.value_or(-1) != -1) {
Flush(BufferOffset);
close(*CodeMapFD);
}
}
bool CodeMapWriter::IsWriteEnabled(const ExecutableFileSectionInfo& Section) {
if (CodeMapFD == -1) {
return false;
}
// PV libraries can't yet be read by FEXServer, so skip dumping them
if (Section.FileInfo.Filename.starts_with("/run/pressure-vessel")) {
return false;
}
if (CodeMapFD) {
return true;
}
// Acquire mutex and re-check CodeMapFD to avoid race conditions
auto lk = std::unique_lock {Mutex};
if (!CodeMapFD) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
return CodeMapFD != -1;
}
void CodeMapWriter::Flush(size_t Offset) {
// Acquire exclusive lock and flush circular buffer
std::unique_lock Lock {Mutex};
Flush(Offset, Lock);
}
void CodeMapWriter::Flush(size_t Offset, std::unique_lock<std::shared_mutex>&) {
write(*CodeMapFD, Buffer.data(), Offset);
BufferOffset = 0;
}
void CodeMapWriter::AppendBlock(const FEXCore::ExecutableFileSectionInfo& SectionInfo, uint64_t BlockEntry) {
if (!IsWriteEnabled(SectionInfo)) {
return;
}
BlockEntry -= SectionInfo.FileStartVA;
if (BlockEntry > std::numeric_limits<uint32_t>::max()) {
ERROR_AND_DIE_FMT("Cannot write code map");
}
// Register new library if not already known
bool NewLibraryLoad = false;
{
// Check prior registration with shared lock
std::shared_lock Lock {Mutex};
NewLibraryLoad = !KnownFileIds.contains(SectionInfo.FileInfo.FileId);
}
if (NewLibraryLoad) {
// Register to map with exclusive lock
std::unique_lock Lock {Mutex};
NewLibraryLoad &= KnownFileIds.insert(SectionInfo.FileInfo.FileId).second;
}
if (NewLibraryLoad) {
// Add entry to code map
AppendLibraryLoad(SectionInfo.FileInfo);
}
// Register the actual code block
CodeMap::Entry DataEntry {SectionInfo.FileInfo.FileId, static_cast<uint32_t>(BlockEntry)};
AppendData(std::as_bytes(std::span {&DataEntry, 1}));
}
void CodeMapWriter::AppendLibraryLoad(const FEXCore::ExecutableFileInfo& FileInfo) {
// See CodeMap::ExternalLibraryInfo
auto ExternalFileId = FileInfo.FileId;
auto TotalSize = AlignUp(sizeof(CodeMap::LoadExternalLibrary) + sizeof(ExternalFileId) + FileInfo.Filename.size() + 1, 4);
const auto Data = reinterpret_cast<char*>(alloca(TotalSize));
auto WritePtr = std::copy_n(reinterpret_cast<const char*>(&CodeMap::LoadExternalLibrary), sizeof(CodeMap::LoadExternalLibrary), Data);
WritePtr = std::copy_n(reinterpret_cast<const char*>(&ExternalFileId), sizeof(ExternalFileId), WritePtr);
WritePtr = std::copy(FileInfo.Filename.begin(), FileInfo.Filename.end(), WritePtr);
std::fill(WritePtr, Data + TotalSize, 0);
AppendData(std::as_bytes(std::span {Data, TotalSize}));
}
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo) {
CodeMap::SetExecutableFileId Data {.ExecutableFileId = FileInfo.FileId};
AppendData(std::span {reinterpret_cast<const std::byte*>(&Data), sizeof(Data)});
}
void CodeMapWriter::AppendData(std::span<const std::byte> Data) {
std::shared_lock Lock {Mutex};
auto Offset = BufferOffset.fetch_add(Data.size_bytes());
if (Offset + Data.size_bytes() > Buffer.size()) {
// Acquire exclusive lock and flush the buffer.
// Under heavy pressure, multiple threads may observe an exhausted buffer simultaneously.
// The thread with the last in-bounds Offset is responsible for flushing the buffer.
Lock.unlock();
bool IsResponsibleForFlush = false;
{
std::unique_lock ExclusiveLock {Mutex};
IsResponsibleForFlush = (Offset <= Buffer.size());
if (IsResponsibleForFlush) {
Flush(Offset, ExclusiveLock);
}
}
if (!IsResponsibleForFlush) {
// Wait for the buffer to be flushed on the responsible thread
Utils::SpinWaitLock::WaitPred<std::less_equal<>, size_t>(reinterpret_cast<size_t*>(&BufferOffset), Buffer.size());
}
AppendData(Data);
return;
}
memcpy(&Buffer.at(Offset), Data.data(), Data.size_bytes());
}
} // namespace FEXCore
namespace FEXCore::Context {
@@ -15,12 +216,156 @@ CodeCache::CodeCache(ContextImpl& CTX_)
: CTX(CTX_) {}
CodeCache::~CodeCache() = default;
uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
if (Filename.empty()) {
return 0xffff'ffff'ffff'ffff;
}
// For now, we just use the file path as an identifier.
// TODO: Ensure the hash is unique enough to distinguish executables while remaining independent of the installation location
return XXH3_64bits(Filename.data(), Filename.size());
}
struct CodeCacheHeader {
char Magic[4] = {'F', 'X', 'C', 'C'};
uint32_t FormatVersion = 1;
char FEXVersion[8] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
uint32_t CodeBufferSize;
uint32_t NumRelocations;
uint64_t SerializedBaseAddress;
// TODO: Consider including information from LookupCache.BlockLinks
};
void CodeCache::LoadData(Core::InternalThreadState& Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& GuestRIPLookup) {
// TODO
}
template<typename T>
static constexpr auto IsOrderedContainer(const T&) -> std::false_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::map<T...>&) -> std::true_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::set<T...>&) -> std::true_type;
bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const ExecutableFileSectionInfo& SourceBinary, uint64_t SerializedBaseAddress) {
// TODO
auto CodeBuffer = CTX.GetLatest();
auto& LookupCache = *Thread.LookupCache->Shared;
auto Relocations = Thread.CPUBackend->TakeRelocations(SourceBinary.FileStartVA);
// Write file header
CodeCacheHeader header;
memcpy(&header.FEXVersion[0], GIT_SHORT_HASH, strlen(GIT_SHORT_HASH));
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
// Dump guest<->host block mappings
{
// Cache contents must be deterministic, so copy the unordered block list and then sort by key
static_assert(!decltype(IsOrderedContainer(LookupCache.BlockList))::value, "Already deterministic; drop temporary container");
fextl::vector<std::pair<uint64_t, const GuestToHostMap::BlockEntry*>> BlockList;
BlockList.reserve(LookupCache.BlockList.size());
for (auto& [Guest, BlockEntry] : LookupCache.BlockList) {
static_assert(sizeof(Guest) == 8, "Breaking change in code cache data layout");
BlockList.emplace_back(Guest, &BlockEntry);
}
std::ranges::sort(BlockList);
for (auto [Guest, Host] : BlockList) {
static_assert(sizeof(Host->HostCode) == 8, "Breaking change in code cache data layout");
static_assert(sizeof(Host->CodePages[0]) == 8, "Breaking change in code cache data layout");
Guest -= SourceBinary.FileStartVA;
::write(fd, &Guest, sizeof(Guest));
uint64_t HostCode = Host->HostCode - reinterpret_cast<uintptr_t>(CodeBuffer->Ptr);
::write(fd, &HostCode, sizeof(HostCode));
uint64_t NumCodePages = Host->CodePages.size();
::write(fd, &NumCodePages, sizeof(NumCodePages));
LOGMAN_THROW_A_FMT(std::ranges::is_sorted(Host->CodePages), "Code pages aren't sorted");
for (auto CodePage : Host->CodePages) {
CodePage -= SourceBinary.FileStartVA;
::write(fd, &CodePage, sizeof(CodePage));
}
}
}
// Dump relocations
static_assert(sizeof(Relocations[0]) == 48, "Breaking change in code cache data layout");
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
}
// Dump the host code (relocated for position-independent serialization)
std::vector CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Dump code pages
static_assert(decltype(IsOrderedContainer(LookupCache.CodePages))::value, "Non-deterministic data source");
for (auto& [Page, Entrypoints] : LookupCache.CodePages) {
static_assert(sizeof(Page) == 8, "Breaking change in code cache data layout");
::write(fd, &Page, sizeof(Page));
uint64_t NumEntrypoints = Entrypoints.size();
::write(fd, &NumEntrypoints, sizeof(NumEntrypoints));
::write(fd, Entrypoints.data(), Entrypoints.size() * sizeof(Entrypoints[0]));
}
return true;
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
// Generate a literal so we can place it
uint64_t Pointer = ForStorage ? 0 : GetNamedSymbolLiteral(CTX, Reloc.NamedSymbolLiteral.Symbol);
Emitter.dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = ForStorage ? 0 : reinterpret_cast<uint64_t>(CTX.ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Emitter.dc64(GuestEntry + Reloc.GuestRIP.GuestRIP);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
uint64_t Pointer = Reloc.GuestRIP.GuestRIP + GuestEntry;
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIP.RegisterIndex), Pointer, true);
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type {}", ToUnderlying(Reloc.Header.Type));
}
}
return true;
}
+47 -38
View File
@@ -437,6 +437,10 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
Profiler::PostForkAction(Child);
if (Child) {
if (CodeMapWriter) {
CodeMapWriter->ResetAfterFork();
}
CodeInvalidationMutex.StealAndDropActiveLocks();
if (Config.StrictInProcessSplitLocks) {
StrictSplitLockMutex = 0;
@@ -459,9 +463,14 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
}
#endif
void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
void ContextImpl::OnCodeBufferAllocated(const fextl::shared_ptr<CPU::CodeBuffer>& Buffer) {
if (Config.GlobalJITNaming()) {
Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
Symbols.RegisterJITSpace(Buffer->Ptr, Buffer->Size);
}
{
std::scoped_lock lk {CodeBufferListLock};
CodeBufferList.emplace_back(Buffer);
}
}
@@ -715,7 +724,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (SourcecodeResolver && Config.GDBSymbols()) {
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (MappedSection) {
MappedSection->FileInfo.SourcecodeMap = SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, MappedSection->FileInfo.FileId);
MappedSection->FileInfo.SourcecodeMap =
SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, CodeMap::GetBaseFilename(MappedSection->FileInfo, false));
}
}
@@ -833,10 +843,14 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
Thread->CPUBackend->ClearRelocations();
}
fextl::vector<uint64_t> CodePages;
if (NeedsAddGuestCodeRanges) {
// Track in the guest to host map all entrypoints for all pages the compiled block touches, if any page didn't previously
// contain code, inform the frontend so it can setup SMC detection.
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
CodePages.reserve(BlockInfo->CodePages.size());
CodePages.insert(CodePages.end(), BlockInfo->CodePages.begin(), BlockInfo->CodePages.end());
for (auto CodePage : BlockInfo->CodePages) {
if (Thread->LookupCache->AddBlockExecutableRange(Thread, BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
@@ -845,8 +859,16 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
// Insert to lookup cache
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, HostAddr);
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, CodePages, HostAddr);
}
if (CodeMapWriter) {
auto Region = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (Region && Region->FileStartVA != 0) {
CodeMapWriter->AppendBlock(*Region, GuestRIP);
}
}
return (uintptr_t)CodePtr;
@@ -873,50 +895,37 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
void ContextImpl::InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) {
FEXCORE_PROFILE_SCOPED("InvalidateCodeBuffersCodeRange");
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::scoped_lock lk {CodeBufferListLock};
auto it = CodeBufferList.begin();
while (it != CodeBufferList.end()) {
if (auto Strong = it->lock()) {
Strong->LookupCache->InvalidateRange(Start, Length);
it++;
} else {
it = CodeBufferList.erase(it);
}
}
}
void ContextImpl::InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
auto lk = Thread->LookupCache->AcquireWriteLock();
auto& CodePages = Thread->LookupCache->Shared->CodePages;
if (Thread->LookupCache->InvalidateCacheRange(Start, Length)) {
FEXCORE_PROFILE_SCOPED("InvalidateCallRet");
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
Accumulator.emplace_back(std::move(it->second));
}
bool InvalidatedAnyEntries = false;
for (const auto& PageEntries : Accumulator) {
for (const auto& Entry : PageEntries) {
if (ContextImpl::ThreadRemoveCodeEntry(Thread, Entry, lk)) {
InvalidatedAnyEntries = true;
}
}
}
if (InvalidatedAnyEntries) {
// This may cause access violations in the thread on Windows as zeroing is not atomic, this is handled by the frontend
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP,
const FEXCore::LookupCacheWriteLockToken& lk) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP, lk);
}
void ContextImpl::ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
static_cast<ContextImpl*>(Frame->Thread->CTX)->SyscallHandler->InvalidateGuestCodeRange(Frame->Thread, GuestRIP, 1);
}
@@ -1,11 +1,10 @@
// SPDX-License-Identifier: MIT
#include "Common/SoftFloat.h"
#include "Common/VectorRegType.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Config/Config.h>
@@ -26,9 +25,7 @@
#endif
#include <array>
#include <atomic>
#include <bit>
#include <condition_variable>
#include <csignal>
#include <cstring>
@@ -95,12 +92,12 @@ void Dispatcher::EmitDispatcher() {
FillStaticRegs();
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
(void)cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
ARMEmitter::BiDirectionalLabel LoopTop {};
#ifdef _M_ARM_64EC
b(&LoopTop);
(void)b(&LoopTop);
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
@@ -108,10 +105,10 @@ void Dispatcher::EmitDispatcher() {
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
// Force a single instruction block if ENTRY_FILL_SRA_SINGLE_INST_REG is nonzero entering the JIT, used for inline SMC handling.
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
(void)cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
// Enter JIT
b(&LoopTop);
(void)b(&LoopTop);
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
// Load ThreadState and write the target PC there
@@ -132,7 +129,7 @@ void Dispatcher::EmitDispatcher() {
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, REG_CALLRET_SP);
// EC_CALL_CHECKER_PC_REG is REG_PF which isn't touched by any of the above
sub(ARMEmitter::Size::i64Bit, TMP1, EC_CALL_CHECKER_PC_REG, TMP1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
// If the entry at the TOS is for the target address, pop it and return to the JIT code
add(ARMEmitter::Size::i64Bit, REG_CALLRET_SP, REG_CALLRET_SP, 0x10);
@@ -144,7 +141,7 @@ void Dispatcher::EmitDispatcher() {
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
Bind(&LoopTop);
(void)Bind(&LoopTop);
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
// Load in our RIP
@@ -171,16 +168,16 @@ void Dispatcher::EmitDispatcher() {
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
Bind(&l_NotECCode);
(void)Bind(&l_NotECCode);
#endif
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbnz(ARMEmitter::Size::i32Bit, TMP1, &CompileSingleStep);
(void)cbnz(ARMEmitter::Size::i32Bit, TMP1, &CompileSingleStep);
ARMEmitter::ForwardLabel NoBlock;
if (DisableL2Cache()) {
b(&NoBlock);
(void)b(&NoBlock);
} else {
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
@@ -203,7 +200,7 @@ void Dispatcher::EmitDispatcher() {
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 3);
// If page pointer is zero then we have no block
cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
// Steal the page offset
and_(ARMEmitter::Size::i64Bit, TMP2, TMP4, 0x0FFF);
@@ -218,10 +215,10 @@ void Dispatcher::EmitDispatcher() {
// If the guest address doesn't match, Compile the block.
sub(TMP2, TMP2, RipReg);
cbnz(ARMEmitter::Size::i64Bit, TMP2, &NoBlock);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP2, &NoBlock);
// Check the host address to see if it matches, else compile the block.
cbz(ARMEmitter::Size::i64Bit, TMP4, &NoBlock);
(void)cbz(ARMEmitter::Size::i64Bit, TMP4, &NoBlock);
// If we've made it here then we have a real compiled block
{
@@ -313,7 +310,7 @@ void Dispatcher::EmitDispatcher() {
// Need to create the block
{
Bind(&NoBlock);
(void)Bind(&NoBlock);
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
@@ -347,7 +344,7 @@ void Dispatcher::EmitDispatcher() {
}
{
Bind(&CompileSingleStep);
(void)Bind(&CompileSingleStep);
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
@@ -491,7 +488,7 @@ void Dispatcher::EmitDispatcher() {
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->X86CodeGen.CallbackReturn);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->SignalDelegation->GetThunkCallbackRET());
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, CTX->Config.Is64BitMode ? 16 : 12);
@@ -509,7 +506,7 @@ void Dispatcher::EmitDispatcher() {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
// Now go back to the regular dispatcher loop
b(&LoopTop);
(void)b(&LoopTop);
}
auto EmitLongALUOpHandler = [&](auto R, auto Offset) {
@@ -578,14 +575,15 @@ void Dispatcher::EmitDispatcher() {
}
}
Bind(&l_CTX);
(void)Bind(&l_CTX);
dc64(reinterpret_cast<uintptr_t>(CTX));
Bind(&l_Sleep);
(void)Bind(&l_Sleep);
dc64(reinterpret_cast<uint64_t>(SleepThread));
Bind(&l_CompileBlock);
(void)Bind(&l_CompileBlock);
FEXCore::Utils::MemberFunctionToPointerCast PMFCompileBlock(&FEXCore::Context::ContextImpl::CompileBlock);
dc64(PMFCompileBlock.GetConvertedPointer());
Bind(&l_CompileSingleStep);
(void)Bind(&l_CompileSingleStep);
FEXCore::Utils::MemberFunctionToPointerCast PMFCompileSingleStep(&FEXCore::Context::ContextImpl::CompileSingleStep);
dc64(PMFCompileSingleStep.GetConvertedPointer());
@@ -51,6 +51,10 @@ public:
}
#endif
uint64_t GetExitFunctionLinkerAddress() const {
return ExitFunctionLinkerAddress;
}
SignalDelegatorConfig MakeSignalDelegatorConfig() const;
protected:
@@ -9,7 +9,6 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/LookupCache.h"
#include <array>
@@ -90,11 +89,6 @@ Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
}
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
if (EntryPoint == CTX->X86CodeGen.CallbackReturn) {
return true;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
+23 -24
View File
@@ -50,14 +50,13 @@ DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
auto Dst = GetReg(Node);
uint64_t Mask = ~0ULL;
const auto OpSize = IROp->Size;
if (OpSize == IR::OpSize::i32Bit) {
Mask = 0xFFFF'FFFFULL;
}
LoadConstant(ARMEmitter::Size::i64Bit, Dst, Constant & Mask);
InsertGuestRIPMove(GetReg(Node), Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -588,7 +587,7 @@ DEF_OP(ShiftFlags) {
and_(ARMEmitter::Size::i32Bit, TMP1, Src2, OpSize == IR::OpSize::i64Bit ? 0x3f : 0x1f);
ARMEmitter::ForwardLabel Done;
cbz(EmitSize, TMP1, &Done);
(void)cbz(EmitSize, TMP1, &Done);
{
// PF/SF/ZF/OF
if (OpSize >= IR::OpSize::i32Bit) {
@@ -652,7 +651,7 @@ DEF_OP(ShiftFlags) {
msr(ARMEmitter::SystemRegister::NZCV, TMP2);
}
}
Bind(&Done);
(void)Bind(&Done);
// TODO: Make RA less dumb so this can't happen (e.g. with late-kill).
if (PFOutput != PFTemp) {
@@ -669,7 +668,7 @@ DEF_OP(RotateFlags) {
// If shift=0, flags are unaffected. Wrap the whole implementation in a cbz.
ARMEmitter::ForwardLabel Done;
cbz(EmitSize, Shift, &Done);
(void)cbz(EmitSize, Shift, &Done);
{
// Extract the last bit shifted in to CF
const auto BitSize = IR::OpSizeToSize(Op->Size) * 8;
@@ -701,7 +700,7 @@ DEF_OP(RotateFlags) {
msr(ARMEmitter::SystemRegister::NZCV, TMP3);
}
}
Bind(&Done);
(void)Bind(&Done);
}
DEF_OP(Extr) {
@@ -767,14 +766,14 @@ DEF_OP(PDep) {
// Now, they're copied, so we can start setting Dest (even if it overlaps with
// one of them). Handle early exit case
mov(EmitSize, Dest, 0);
cbz(EmitSize, OrigMask, &Done);
(void)cbz(EmitSize, OrigMask, &Done);
// Setup for first iteration
neg(EmitSize, T0, Mask);
and_(EmitSize, T0, T0, Mask);
// Main loop
Bind(&NextBit);
(void)Bind(&NextBit);
sbfx(EmitSize, T1, Input, 0, 1);
eor(EmitSize, Mask, Mask, T0);
and_(EmitSize, T0, T1, T0);
@@ -782,10 +781,10 @@ DEF_OP(PDep) {
orr(EmitSize, Dest, Dest, T0);
lsr(EmitSize, Input, Input, 1);
and_(EmitSize, T0, Mask, T1);
cbnz(EmitSize, T0, &NextBit);
(void)cbnz(EmitSize, T0, &NextBit);
// All done with nothing to do.
Bind(&Done);
(void)Bind(&Done);
}
}
@@ -821,27 +820,27 @@ DEF_OP(PExt) {
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::ForwardLabel Done;
cbz(EmitSize, Mask, &EarlyExit);
(void)cbz(EmitSize, Mask, &EarlyExit);
mov(EmitSize, MaskReg, Mask);
mov(EmitSize, ValueReg, Input);
mov(EmitSize, Dest, ARMEmitter::Reg::zr);
// Main loop
Bind(&NextBit);
cbz(EmitSize, MaskReg, &Done);
(void)Bind(&NextBit);
(void)cbz(EmitSize, MaskReg, &Done);
clz(EmitSize, BitReg, MaskReg);
lslv(EmitSize, ValueReg, ValueReg, BitReg);
lslv(EmitSize, MaskReg, MaskReg, BitReg);
extr(EmitSize, Dest, Dest, ValueReg, OpSizeBitsM1);
bfc(EmitSize, MaskReg, OpSizeBitsM1, 1);
b(&NextBit);
(void)b(&NextBit);
// Early exit
Bind(&EarlyExit);
(void)Bind(&EarlyExit);
mov(EmitSize, Dest, ARMEmitter::Reg::zr);
// All done with nothing to do.
Bind(&Done);
(void)Bind(&Done);
}
}
@@ -909,7 +908,7 @@ DEF_OP(Div) {
eor(EmitSize, TMP1, TMP1, Upper);
// If the sign bit matches then the result is zero
cbz(EmitSize, TMP1, &Only64Bit);
(void)cbz(EmitSize, TMP1, &Only64Bit);
// Long divide
{
@@ -928,17 +927,17 @@ DEF_OP(Div) {
mov(EmitSize, Remainder, TMP2);
// Skip 64-bit path
b(&LongDIVRet);
(void)b(&LongDIVRet);
}
Bind(&Only64Bit);
(void)Bind(&Only64Bit);
// 64-Bit only
{
sdiv(EmitSize, Quotient, Lower, Divisor);
msub(EmitSize, Remainder, Quotient, Divisor, Lower);
}
Bind(&LongDIVRet);
(void)Bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown DIV Size: {}", OpSize); break;
@@ -992,7 +991,7 @@ DEF_OP(UDiv) {
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
cbz(EmitSize, Upper, &Only64Bit);
(void)cbz(EmitSize, Upper, &Only64Bit);
// Long divide
{
@@ -1011,17 +1010,17 @@ DEF_OP(UDiv) {
mov(EmitSize, Remainder, TMP2);
// Skip 64-bit path
b(&LongDIVRet);
(void)b(&LongDIVRet);
}
Bind(&Only64Bit);
(void)Bind(&Only64Bit);
// 64-Bit only
{
udiv(EmitSize, Quotient, Lower, Divisor);
msub(EmitSize, Remainder, Quotient, Divisor, Lower);
}
Bind(&LongDIVRet);
(void)Bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LUDIV Size: {}", OpSize); break;
@@ -11,23 +11,18 @@ $end_info$
#include <FEXCore/Core/Thunks.h>
namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl& CTX, FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op)); break;
return CTX.Dispatcher->GetExitFunctionLinkerAddress();
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum& Sum) {
Relocation MoveABI {};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE};
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
@@ -38,9 +33,9 @@ void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR
}
Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
uint64_t Pointer = GetNamedSymbolLiteral(*CTX, Op);
Arm64JITCore::NamedSymbolLiteralPair Lit {
NamedSymbolLiteralPair Lit {
.Lit = Pointer,
.MoveABI =
{
@@ -48,92 +43,72 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
{
.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit) {
switch (Lit.MoveABI.Header.Type) {
case RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL:
case RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Lit.MoveABI.Header.Offset = GetCursorOffset();
break;
}
Bind(&Lit.Loc);
default: ERROR_AND_DIE_FMT("Unknown relocation type for {}", __FUNCTION__);
}
BindOrRestart(&Lit.Loc);
dc64(Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
auto Arm64JITCore::InsertGuestRIPLiteral(uint64_t GuestRIP) -> NamedSymbolLiteralPair {
return {
.Lit = GuestRIP,
.MoveABI =
{
.GuestRIP = {.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL,
},
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
.GuestRIP = GuestRIP},
},
};
}
void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant) {
Relocation MoveABI {};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
MoveABI.GuestRIP.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE};
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
MoveABI.GuestRIP.GuestRIP = Constant;
MoveABI.GuestRIP.RegisterIndex = Reg.Idx();
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, false);
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation> Relocations) {
const auto OrigBase = GetBufferBase();
const auto OrigSize = GetBufferSize();
const auto OrigOffset = GetCursorOffset();
SetBuffer(reinterpret_cast<std::uint8_t*>(Code.data()), Code.size_bytes());
for (auto& Reloc : Relocations) {
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc.NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(Reloc.NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIPMove.RegisterIndex), Pointer, true);
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations(uint64_t GuestBaseAddress) {
// Rebase relocations to library base address
for (auto& Relocation : Relocations) {
switch (Relocation.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE:
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Relocation.GuestRIP.GuestRIP -= GuestBaseAddress;
break;
}
default:;
}
}
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return true;
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations() {
return std::move(Relocations);
}
+32 -52
View File
@@ -62,27 +62,27 @@ DEF_OP(CASPair) {
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::ForwardLabel LoopNotExpected;
ARMEmitter::ForwardLabel LoopExpected;
Bind(&LoopTop);
(void)Bind(&LoopTop);
// This instruction sequence must be synced with HandleCASPAL_Armv8.
ldaxp(EmitSize, TMP2, TMP3, MemSrc);
cmp(EmitSize, TMP2, Expected0);
ccmp(EmitSize, TMP3, Expected1, ARMEmitter::StatusFlags::None, ARMEmitter::Condition::CC_EQ);
b(ARMEmitter::Condition::CC_NE, &LoopNotExpected);
(void)b(ARMEmitter::Condition::CC_NE, &LoopNotExpected);
stlxp(EmitSize, TMP2, Desired0, Desired1, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
(void)cbnz(EmitSize, TMP2, &LoopTop);
mov(EmitSize, Dst0, Expected0);
mov(EmitSize, Dst1, Expected1);
b(&LoopExpected);
(void)b(&LoopExpected);
Bind(&LoopNotExpected);
(void)Bind(&LoopNotExpected);
mov(EmitSize, Dst0, TMP2.R());
mov(EmitSize, Dst1, TMP3.R());
// exclusive monitor needs to be cleared here
// Might have hit the case where ldaxr was hit but stlxr wasn't
clrex();
Bind(&LoopExpected);
(void)Bind(&LoopExpected);
// Restore
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
@@ -114,7 +114,7 @@ DEF_OP(CAS) {
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::ForwardLabel LoopNotExpected;
ARMEmitter::ForwardLabel LoopExpected;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
if (IROp->Size == IR::OpSize::i8Bit) {
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTB, 0);
@@ -123,38 +123,18 @@ DEF_OP(CAS) {
} else {
cmp(EmitSize, TMP2, Expected);
}
b(ARMEmitter::Condition::CC_NE, &LoopNotExpected);
(void)b(ARMEmitter::Condition::CC_NE, &LoopNotExpected);
stlxr(SubEmitSize, TMP3, Desired, MemSrc);
cbnz(EmitSize, TMP3, &LoopTop);
(void)cbnz(EmitSize, TMP3, &LoopTop);
mov(EmitSize, Dst, Expected);
b(&LoopExpected);
(void)b(&LoopExpected);
Bind(&LoopNotExpected);
(void)Bind(&LoopNotExpected);
mov(EmitSize, Dst, TMP2.R());
// exclusive monitor needs to be cleared here
// Might have hit the case where ldaxr was hit but stlxr wasn't
clrex();
Bind(&LoopExpected);
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
eor(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
(void)Bind(&LoopExpected);
}
}
@@ -179,10 +159,10 @@ DEF_OP(AtomicSwap) {
ldswpal(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
stlxr(SubEmitSize, TMP4, Src, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
ubfm(EmitSize, GetReg(Node), TMP2, 0, IR::OpSizeAsBits(OpSize) - 1);
}
}
@@ -199,11 +179,11 @@ DEF_OP(AtomicFetchAdd) {
ldaddal(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
add(EmitSize, TMP3, TMP2, Src);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
@@ -221,11 +201,11 @@ DEF_OP(AtomicFetchSub) {
ldaddal(SubEmitSize, TMP2, GetReg(Node), MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
sub(EmitSize, TMP3, TMP2, Src);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
@@ -243,11 +223,11 @@ DEF_OP(AtomicFetchAnd) {
ldclral(SubEmitSize, TMP2, GetReg(Node), MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
and_(EmitSize, TMP3, TMP2, Src);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
@@ -264,11 +244,11 @@ DEF_OP(AtomicFetchCLR) {
ldclral(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
bic(EmitSize, TMP3, TMP2, Src);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
@@ -285,11 +265,11 @@ DEF_OP(AtomicFetchOr) {
ldsetal(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
orr(EmitSize, TMP3, TMP2, Src);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
@@ -306,11 +286,11 @@ DEF_OP(AtomicFetchXor) {
ldeoral(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
eor(EmitSize, TMP3, TMP2, Src);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
@@ -326,20 +306,20 @@ DEF_OP(AtomicFetchNeg) {
// Use a CAS loop to avoid needing to emulate unaligned LLSC atomics
ldr(SubEmitSize, TMP2, MemSrc);
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
mov(EmitSize, TMP4, TMP2);
neg(EmitSize, TMP3, TMP2);
casal(SubEmitSize, TMP2, TMP3, MemSrc);
sub(EmitSize, TMP3, TMP2, TMP4);
cbnz(EmitSize, TMP3, &LoopTop);
(void)cbnz(EmitSize, TMP3, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
neg(EmitSize, TMP3, TMP2);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
(void)cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
@@ -359,11 +339,11 @@ DEF_OP(TelemetrySetValue) {
stsetl(ARMEmitter::SubRegSize::i64Bit, TMP1, TMP2);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
(void)Bind(&LoopTop);
ldaxr(ARMEmitter::SubRegSize::i64Bit, TMP3, TMP2);
orr(ARMEmitter::Size::i32Bit, TMP3, TMP3, Src);
stlxr(ARMEmitter::SubRegSize::i64Bit, TMP3, TMP3, TMP2);
cbnz(ARMEmitter::Size::i32Bit, TMP3, &LoopTop);
(void)cbnz(ARMEmitter::Size::i32Bit, TMP3, &LoopTop);
}
#endif
}
+19 -20
View File
@@ -141,7 +141,7 @@ DEF_OP(ExitFunction) {
if (!Op->CallReturnBlock.IsInvalid()) {
auto CallReturnAddressReg = GetReg(Op->CallReturnAddress).X();
PendingCallReturnTargetLabel = &CallReturnTargets.try_emplace(Op->CallReturnBlock.ID()).first->second;
adr(TMP1, &l_CallReturn);
(void)adr(TMP1, &l_CallReturn);
stp<ARMEmitter::IndexType::PRE>(CallReturnAddressReg, TMP1, REG_CALLRET_SP, -0x10);
} else {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
@@ -149,16 +149,16 @@ DEF_OP(ExitFunction) {
} else if (Op->Hint == IR::BranchHint::CheckTF) {
ARMEmitter::ForwardLabel TFUnset;
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbz(ARMEmitter::Size::i32Bit, TMP1, &TFUnset);
(void)cbz(ARMEmitter::Size::i32Bit, TMP1, &TFUnset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, NewRIP);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
blr(TMP2);
Bind(&TFUnset);
(void)Bind(&TFUnset);
}
EmitLinkedBranch(NewRIP, Op->Hint == IR::BranchHint::Call);
Bind(&l_CallReturn);
(void)Bind(&l_CallReturn);
#ifdef _M_ARM_64EC
}
#endif
@@ -170,7 +170,7 @@ DEF_OP(ExitFunction) {
// First try to pop from the call-ret stack, otherwise follow the normal path (but ending in a ret)
ldp<ARMEmitter::IndexType::POST>(TMP1, TMP2, REG_CALLRET_SP, 0x10);
sub(TMP1, TMP1, RipReg.X());
cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
}
// L1 Cache
@@ -185,23 +185,23 @@ DEF_OP(ExitFunction) {
// Note: sub+cbnz used over cmp+br to preserve flags.
sub(TMP1, TMP1, RipReg.X());
cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
Bind(&SkipFullLookup);
(void)Bind(&SkipFullLookup);
if (Op->Hint == IR::BranchHint::Call) {
ARMEmitter::ForwardLabel l_CallReturn;
if (!Op->CallReturnBlock.IsInvalid()) {
auto CallReturnAddressReg = GetReg(Op->CallReturnAddress).X();
PendingCallReturnTargetLabel = &CallReturnTargets.try_emplace(Op->CallReturnBlock.ID()).first->second;
adr(TMP1, &l_CallReturn);
(void)adr(TMP1, &l_CallReturn);
stp<ARMEmitter::IndexType::PRE>(CallReturnAddressReg, TMP1, REG_CALLRET_SP, -0x10);
} else {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
}
blr(TMP2);
Bind(&l_CallReturn);
(void)Bind(&l_CallReturn);
} else if (Op->Hint == IR::BranchHint::Return) {
ret(TMP2);
} else {
@@ -222,7 +222,7 @@ DEF_OP(CondJump) {
auto TrueTargetLabel = JumpTarget(Op->TrueBlock);
if (Op->FromNZCV) {
b(MapCC(Op->Cond), TrueTargetLabel);
b_OrRestart(MapCC(Op->Cond), TrueTargetLabel);
} else {
uint64_t Const;
const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
@@ -235,16 +235,16 @@ DEF_OP(CondJump) {
if (Op->Cond == IR::CondClass::EQ) {
LOGMAN_THROW_A_FMT(Const == 0, "CondJump: Expected 0 source");
cbz(Size, Reg, TrueTargetLabel);
cbz_OrRestart(Size, Reg, TrueTargetLabel);
} else if (Op->Cond == IR::CondClass::NEQ) {
LOGMAN_THROW_A_FMT(Const == 0, "CondJump: Expected 0 source");
cbnz(Size, Reg, TrueTargetLabel);
cbnz_OrRestart(Size, Reg, TrueTargetLabel);
} else if (Op->Cond == IR::CondClass::TSTZ) {
LOGMAN_THROW_A_FMT(Const < 64, "CondJump: Expected valid bit source");
tbz(Reg, Const, TrueTargetLabel);
tbz_OrRestart(Reg, Const, TrueTargetLabel);
} else if (Op->Cond == IR::CondClass::TSTNZ) {
LOGMAN_THROW_A_FMT(Const < 64, "CondJump: Expected valid bit source");
tbnz(Reg, Const, TrueTargetLabel);
tbnz_OrRestart(Reg, Const, TrueTargetLabel);
} else {
LOGMAN_THROW_A_FMT(false, "CondJump expected simple condition");
}
@@ -328,8 +328,7 @@ DEF_OP(Thunk) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr));
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
InsertNamedThunkRelocation(ARMEmitter::Reg::r2, Op->ThunkNameHash);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
} else {
@@ -355,7 +354,7 @@ DEF_OP(ValidateCode) {
while (len >= Size) {
LoadData();
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &Fail);
cbnz_OrRestart(ARMEmitter::Size::i64Bit, TMP1, &Fail);
len -= Size;
Offset += Size;
}
@@ -383,10 +382,10 @@ DEF_OP(ValidateCode) {
ARMEmitter::ForwardLabel End;
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 0);
b(&End);
Bind(&Fail);
b_OrRestart(&End);
BindOrRestart(&Fail);
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 1);
Bind(&End);
BindOrRestart(&End);
}
DEF_OP(ThreadRemoveCodeEntry) {
+95 -67
View File
@@ -11,8 +11,6 @@ desc: Main glue logic of the arm64 splatter backend
$end_info$
*/
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
@@ -30,6 +28,7 @@ $end_info$
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/Utils/TypeDefines.h>
@@ -37,7 +36,6 @@ $end_info$
#include <cstdio>
#include <cstring>
#include <limits>
#include <unistd.h>
namespace {
@@ -495,7 +493,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
static void DirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = JumpThunkStartAddress / 4 - CallerAddress / 4;
@@ -513,11 +511,12 @@ static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Co
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
static void IndirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
BranchEmit.b(0x8);
// Restore branch +2 instructions to jump to the linker block
BranchEmit.b(0x2);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
@@ -579,16 +578,11 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
if (KnownCallMarkerInst == ExpectedKnownCallMarkerInst) {
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Frame, Record, true); }, lk);
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, true); }, lk);
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, false);
},
lk);
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, false); }, lk);
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
@@ -742,11 +736,11 @@ void Arm64JITCore::EmitTFCheck() {
// Note that this needs to be before the below suspend checks, as X86 checks this flag immediately after executing an instruction.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbz(ARMEmitter::Size::i32Bit, TMP1, &l_TFUnset);
(void)cbz(ARMEmitter::Size::i32Bit, TMP1, &l_TFUnset);
// X86 semantically checks TF after executing each instruction, so e.g. setting a context with TF set will execute a single instruction
// and then raise an exception. However on the FEX side this is simpler to implement by checking at the start of each instruction, handle this by having bit 1 being unset in the flag state indicate that TF is blocked for a single instruction.
tbz(TMP1, 1, &l_TFBlocked);
(void)tbz(TMP1, 1, &l_TFBlocked);
// Block TF for a single instruction when the frontend jumps to a new context by unsetting bit 1.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
@@ -769,11 +763,11 @@ void Arm64JITCore::EmitTFCheck() {
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
Bind(&l_TFBlocked);
(void)Bind(&l_TFBlocked);
// If TF was blocked for this instruction, unblock it for the next.
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 0b11);
strb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
Bind(&l_TFUnset);
(void)Bind(&l_TFUnset);
}
void Arm64JITCore::EmitSuspendInterruptCheck() {
@@ -789,16 +783,16 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
ldr(TMP2.W(), STATE_PTR(CpuStateFrame, SuspendDoorbell));
ARMEmitter::ForwardLabel l_NoSuspend;
cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
(void)cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
brk(SuspendMagic);
Bind(&l_NoSuspend);
(void)Bind(&l_NoSuspend);
#endif
}
void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool CheckTF) {
// Get the address of the JITCodeHeader and store in to the core state.
// Two instruction cost, each 1 cycle.
adr(TMP1, &HeaderLabel);
adr_OrRestart(TMP1, &HeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
if (CheckTF) {
@@ -815,36 +809,65 @@ void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool C
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
EmitSuspendInterruptCheck();
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
JumpTargets.clear();
CallReturnTargets.clear();
PendingJumpThunks.clear();
uint32_t SSACount = IR->GetSSACount();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
const auto PrevNumAllocations = Relocations.size();
this->Entry = Entry;
this->DebugData = DebugData;
this->IR = IR;
RequiresFarARM64Jumps = false;
SSANodeMultiplier = 24;
// Prepare restart via long jump in case branch encoding fails.
// This uses UncheckedLongJump since we don't implement std::longjmp in WoA setups
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(ThreadState->RestartJump))) {
case RestartOptions::Control::Incoming:
// Nothing
break;
case RestartOptions::Control::EnableFarARM64Jumps: RequiresFarARM64Jumps = true; break;
case RestartOptions::Control::NeedsLargerJITSpace:
// Get rid of the claimed buffer immediately, we can't fit in it at all.
TempAllocator.UnclaimBuffer();
SSANodeMultiplier *= 2;
break;
default: LOGMAN_MSG_A_FMT("Unhandled Arm64 restart condition!");
}
uint32_t SSACount = IR->GetSSACount();
JumpTargets.clear();
CallReturnTargets.clear();
PendingJumpThunks.clear();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
CodeData.EntryPoints.clear();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = 0x1000 + SSACount * 24;
// One page baseline, plus SSANodeMultipler bytes, plus another page for guard page.
const uint32_t DesiredBufferRange = AlignUp(FEXCore::Utils::FEX_PAGE_SIZE * 2 + SSACount * SSANodeMultiplier, FEXCore::Utils::FEX_PAGE_SIZE);
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBuffer = TempAllocator.ReownOrClaimBuffer(BufferRange);
SetBuffer(TempCodeBuffer, BufferRange);
auto TempCodeBufferInfo = TempAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBuffer = TempCodeBufferInfo.Ptr;
const uint32_t UsableBufferRange = TempCodeBufferInfo.Size - FEXCore::Utils::FEX_PAGE_SIZE;
SetBuffer(TempCodeBuffer, UsableBufferRange);
ThreadState->JITGuardPage = reinterpret_cast<uintptr_t>(TempCodeBuffer) + UsableBufferRange;
ThreadState->JITGuardOverflowArgument = FEXCore::ToUnderlying(RestartOptions::Control::NeedsLargerJITSpace);
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
// Put the code header at the start of the data block.
ARMEmitter::BackwardLabel JITCodeHeaderLabel {};
Bind(&JITCodeHeaderLabel);
(void)Bind(&JITCodeHeaderLabel);
JITCodeHeader* CodeHeader = GetCursorAddress<JITCodeHeader*>();
CursorIncrement(sizeof(JITCodeHeader));
@@ -892,7 +915,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
if (PendingTargetLabel->Backward.Location) {
EmitSuspendInterruptCheck();
}
b(PendingTargetLabel);
b_OrRestart(PendingTargetLabel);
PendingTargetLabel = nullptr;
}
@@ -902,14 +925,14 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
const auto IsReturnTarget = CallReturnTargets.try_emplace(Node).first;
if (PendingTargetLabel) {
// If there is a fallthrough branch to this block, skip over the entrypoint code.
b(Target);
b_OrRestart(Target);
} else if (PendingCallReturnTargetLabel && PendingCallReturnTargetLabel != &IsReturnTarget->second) {
// If we just emitted a call, but the block we're now emitting is not the return block so don't fallthrough.
b(PendingCallReturnTargetLabel);
b_OrRestart(PendingCallReturnTargetLabel);
}
PendingCallReturnTargetLabel = nullptr;
Bind(&IsReturnTarget->second);
BindOrRestart(&IsReturnTarget->second);
CodeData.EntryPoints.emplace(BlockStartRIP, GetCursorAddress<uint8_t*>());
DebugData->GuestOpcodes.push_back({BlockIROp->GuestEntryOffset, GetCursorAddress<uint8_t*>() - CodeData.BlockBegin});
@@ -918,12 +941,12 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
if (PendingCallReturnTargetLabel) {
// If there is still a pending call return target, then the block we're emitting is not the return block so don't fallthrough.
b(PendingCallReturnTargetLabel);
b_OrRestart(PendingCallReturnTargetLabel);
PendingCallReturnTargetLabel = nullptr;
}
PendingTargetLabel = nullptr;
Bind(Target);
BindOrRestart(Target);
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
@@ -948,7 +971,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
if (PendingTargetLabel->Backward.Location) {
EmitSuspendInterruptCheck();
}
b(PendingTargetLabel);
b_OrRestart(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
@@ -959,31 +982,37 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
ARMEmitter::ForwardLabel l_DoLink;
uint64_t ThunkAddress = GetCursorAddress<uint64_t>();
Bind(&PendingJumpThunk.Label);
b(&l_DoLink);
BindOrRestart(&PendingJumpThunk.Label);
b_OrRestart(&l_DoLink);
br(TMP1);
Bind(&l_DoLink);
BindOrRestart(&l_DoLink);
ldr(TMP1, &l_ExitLink);
blr(TMP1);
// This is a ExitFunctionLinkData struct
Bind(&l_ExitLink);
dc64(0); // HostCode
dc64(PendingJumpThunk.GuestRIP); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
BindOrRestart(&l_ExitLink);
dc64(0); // HostCode
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(PendingJumpThunk.GuestRIP)); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
Bind(&l_ExitLink);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
BindOrRestart(&l_ExitLink);
PlaceNamedSymbolLiteral(InsertNamedSymbolLiteral(RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER));
// CodeSize not including the header or tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeBegin;
// Add the JitCodeTail
// Add the JitCodeTail (written later)
Align(alignof(JITCodeTail));
auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
const auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
JITCodeTail JITBlockTail {
.RIP = Entry,
.GuestSize = Size,
.SpinLockFutex = 0,
.SingleInst = SingleInst,
};
// Entries that live after the JITCodeTail.
// These entries correlate JIT code regions with guest RIP regions.
@@ -1001,23 +1030,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// FEXCore::Utils::vl64 GuestRIPOffset;
// };
auto JITRIPEntriesBegin = GetCursorAddress<uint8_t*>();
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
JITBlockTail->GuestSize = Size;
JITBlockTail->SingleInst = SingleInst;
JITBlockTail->SpinLockFutex = 0;
const auto JITRIPEntriesBegin = JITBlockTailLocation + sizeof(JITBlockTail);
auto JITRIPEntriesLocation = JITRIPEntriesBegin;
{
// Store the RIP entries.
JITBlockTail->NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail->OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
JITBlockTail.NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail.OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
uintptr_t CurrentRIPOffset = 0;
uint64_t CurrentPCOffset = 0;
@@ -1033,14 +1052,20 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
}
CursorIncrement(JITRIPEntriesLocation - JITRIPEntriesBegin);
SetCursorOffset(JITRIPEntriesLocation - CodeData.BlockBegin);
Align();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
// Finalize and write block tail data
JITBlockTail.Size = CodeData.Size;
{
auto PrevCur = GetCursorOffset();
memcpy(JITBlockTailLocation, &JITBlockTail, sizeof(JITBlockTail));
SetCursorOffset(JITBlockTailLocation - CodeData.BlockBegin + offsetof(JITCodeTail, RIP));
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(JITBlockTail.RIP));
SetCursorOffset(PrevCur);
}
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
@@ -1049,7 +1074,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Query size of generated code
const auto TempSize = GetCursorOffset();
LOGMAN_THROW_A_FMT(TempSize <= BufferRange, "Exceeded bounds of temporary buffer ({:#x} vs {:#x})", TempSize, BufferRange);
// Bring CodeBuffer up to date
{
@@ -1081,6 +1105,10 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
CodeBegin += Delta;
for (std::size_t Idx = PrevNumAllocations; Idx != Relocations.size(); ++Idx) {
Relocations[Idx].Header.Offset += CodeBuffers.LatestOffset;
}
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
+208 -10
View File
@@ -23,6 +23,7 @@ $end_info$
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/LongJump.h>
#include <CodeEmitter/Emitter.h>
@@ -66,6 +67,21 @@ private:
const bool HostSupportsRPRES {};
const bool HostSupportsAFP {};
struct RestartOptions {
enum class Control : uint64_t {
Incoming = 0,
EnableFarARM64Jumps = 1,
NeedsLargerJITSpace = 2,
};
};
// FEXCore makes assumptions in the JIT about certain conditions being true.
// In the rare case when those assumptions are broken, FEX needs to safely restart the JIT.
RestartOptions RestartControl {};
bool RequiresFarARM64Jumps {};
// Default to 6 instructions per SSA node.
uint32_t SSANodeMultiplier {24};
ARMEmitter::BiDirectionalLabel* PendingTargetLabel {};
ARMEmitter::BiDirectionalLabel* PendingCallReturnTargetLabel {};
FEXCore::Context::ContextImpl* CTX {};
@@ -329,14 +345,187 @@ private:
void EmitLinkedBranch(uint64_t GuestRIP, bool Call) {
PendingJumpThunks.push_back({GetCursorAddress<uint64_t>(), GuestRIP, {}});
auto& Thunk = PendingJumpThunks.back();
Bind(&Thunk.Label);
BindOrRestart(&Thunk.Label);
if (Call) {
bl(&Thunk.Label);
bl_OrRestart(&Thunk.Label);
} else {
b(&Thunk.Label);
b_OrRestart(&Thunk.Label);
}
}
// Restart helpers
template<ARMEmitter::IsLabel T>
void bl_OrRestart(T* Label) {
if (bl(Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void b_OrRestart(T* Label) {
if (b(Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void b_OrRestart(ARMEmitter::Condition Cond, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
// Wrap a manual Cond check around an unconditional branch; this can encode larger offsets
(void)b(InvertCondition(Cond), &Skip);
if (b(Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
}
(void)Bind(&Skip);
return;
}
if (b(Cond, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void cbz_OrRestart(ARMEmitter::Size s, ARMEmitter::Register rt, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
// Wrap a manual Cond check around an unconditional branch; this can encode larger offsets
(void)cbnz(s, rt, &Skip);
if (b(Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
}
(void)Bind(&Skip);
return;
}
if (cbz(s, rt, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void cbnz_OrRestart(ARMEmitter::Size s, ARMEmitter::Register rt, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
// Wrap a manual Cond check around an unconditional branch; this can encode larger offsets
(void)cbz(s, rt, &Skip);
if (b(Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
}
(void)Bind(&Skip);
return;
}
if (cbnz(s, rt, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void tbz_OrRestart(ARMEmitter::Register rt, uint32_t Bit, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
// Wrap a manual Cond check around an unconditional branch; this can encode larger offsets
(void)tbnz(rt, Bit, &Skip);
if (b(Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
}
(void)Bind(&Skip);
return;
}
if (tbz(rt, Bit, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void tbnz_OrRestart(ARMEmitter::Register rt, uint32_t Bit, T* Label) {
if (RequiresFarARM64Jumps) {
ARMEmitter::ForwardLabel Skip {};
// Wrap a manual Cond check around an unconditional branch; this can encode larger offsets
(void)tbz(rt, Bit, &Skip);
if (b(Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
}
(void)Bind(&Skip);
return;
}
if (tbnz(rt, Bit, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adr_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADR.");
}
return;
}
if (adr(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adrp_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADRP.");
}
return;
}
if (adrp(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void BindOrRestart(T* Label) {
if (Bind(Label)) {
return;
}
if (RequiresFarARM64Jumps) {
// This should have been caught before this point.
ERROR_AND_DIE_FMT("Unhandled long bind");
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass* RAPass {};
@@ -347,8 +536,6 @@ private:
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
@@ -385,19 +572,30 @@ private:
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Inserts a relocation for a constant value relative to the guest entrypoint
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
NamedSymbolLiteralPair InsertGuestRIPLiteral(uint64_t GuestRIP);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit);
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit);
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation>);
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() override;
/**
* Returns any relocations generated since the last call to TakeRelocations.
*
* GuestBaseAddress must match the base virtual address to which the
* input x86 binary is mapped.
*/
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) override;
/** @} */
+47 -47
View File
@@ -912,7 +912,7 @@ DEF_OP(VLoadVectorMasked) {
// If the sign bit is zero then skip the load
ARMEmitter::ForwardLabel Skip {};
tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
(void)tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
// Do the gather load for this element into the destination
switch (IROp->ElementSize) {
case IR::OpSize::i8Bit: ld1<ARMEmitter::SubRegSize::i8Bit>(TempDst.Q(), i, TempMemReg); break;
@@ -923,7 +923,7 @@ DEF_OP(VLoadVectorMasked) {
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, IROp->ElementSize); return;
}
Bind(&Skip);
(void)Bind(&Skip);
if ((i + 1) != NumElements) {
// Handle register rename to save a move.
@@ -1013,7 +1013,7 @@ DEF_OP(VStoreVectorMasked) {
// If the sign bit is zero then skip the load
ARMEmitter::ForwardLabel Skip {};
tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
(void)tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
// Do the gather load for this element into the destination
switch (IROp->ElementSize) {
case IR::OpSize::i8Bit: st1<ARMEmitter::SubRegSize::i8Bit>(RegData.Q(), i, TempMemReg); break;
@@ -1024,7 +1024,7 @@ DEF_OP(VStoreVectorMasked) {
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, IROp->ElementSize); return;
}
Bind(&Skip);
(void)Bind(&Skip);
if ((i + 1) != NumElements) {
// Handle register rename to save a move.
@@ -1102,7 +1102,7 @@ void Arm64JITCore::Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize,
PerformMove(ElementSize, WorkingReg, MaskReg, i);
// Skip if the mask's sign bit isn't set
tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
(void)tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
// Extract Index Element
if ((IndexElement * IR::OpSizeToSize(VectorIndexSize)) >= 16) {
@@ -1140,7 +1140,7 @@ void Arm64JITCore::Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize,
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, ElementSize); FEX_UNREACHABLE;
}
Bind(&Skip);
(void)Bind(&Skip);
}
if (NeedsDestTmp) {
@@ -1874,7 +1874,7 @@ DEF_OP(MemSet) {
if (!DirectionIsInline) {
// Backward or forwards implementation depends on flag
tbnz(DirectionReg, 1, &BackwardImpl);
(void)tbnz(DirectionReg, 1, &BackwardImpl);
}
auto MemStore = [this](auto Value, uint32_t OpSize, int32_t Size) {
@@ -1922,7 +1922,7 @@ DEF_OP(MemSet) {
ARMEmitter::ForwardLabel DoneInternal {};
// Early exit if zero count.
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
ARMEmitter::ForwardLabel AgainInternal256Exit {};
@@ -1939,50 +1939,50 @@ DEF_OP(MemSet) {
// Do this in two parts, to fallback to the byte by byte loop if size < 32, and to the
// single copy loop if size < 64.
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbnz(TMP1, 63, &AgainInternal128Exit);
(void)tbnz(TMP1, 63, &AgainInternal128Exit);
// Fill VTMP2 with the set pattern
dup(SubRegSize, VTMP2.Q(), Value);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbnz(TMP1, 63, &AgainInternal256Exit);
(void)tbnz(TMP1, 63, &AgainInternal256Exit);
Bind(&AgainInternal256);
(void)Bind(&AgainInternal256);
stp<ARMEmitter::IndexType::POST>(VTMP2.Q(), VTMP2.Q(), TMP2, 32 * Direction);
stp<ARMEmitter::IndexType::POST>(VTMP2.Q(), VTMP2.Q(), TMP2, 32 * Direction);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 64 / Size);
tbz(TMP1, 63, &AgainInternal256);
(void)tbz(TMP1, 63, &AgainInternal256);
Bind(&AgainInternal256Exit);
(void)Bind(&AgainInternal256Exit);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 64 / Size);
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbnz(TMP1, 63, &AgainInternal128Exit);
Bind(&AgainInternal128);
(void)tbnz(TMP1, 63, &AgainInternal128Exit);
(void)Bind(&AgainInternal128);
stp<ARMEmitter::IndexType::POST>(VTMP2.Q(), VTMP2.Q(), TMP2, 32 * Direction);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbz(TMP1, 63, &AgainInternal128);
(void)tbz(TMP1, 63, &AgainInternal128);
Bind(&AgainInternal128Exit);
(void)Bind(&AgainInternal128Exit);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
}
Bind(&AgainInternal);
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemStoreTSO(Value, OpSize, SizeDirection);
} else {
MemStore(Value, OpSize, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
Bind(&DoneInternal);
(void)Bind(&DoneInternal);
if (SizeDirection >= 0) {
switch (OpSize) {
@@ -2012,12 +2012,12 @@ DEF_OP(MemSet) {
EmitMemset(Direction);
if (Direction == 1) {
b(&Done);
Bind(&BackwardImpl);
(void)b(&Done);
(void)Bind(&BackwardImpl);
}
}
Bind(&Done);
(void)Bind(&Done);
// Destination already set to the final pointer.
}
}
@@ -2067,7 +2067,7 @@ DEF_OP(MemCpy) {
if (!DirectionIsInline) {
// Backward or forwards implementation depends on flag
tbnz(DirectionReg, 1, &BackwardImpl);
(void)tbnz(DirectionReg, 1, &BackwardImpl);
}
auto MemCpy = [this](uint32_t OpSize, int32_t Size) {
@@ -2164,7 +2164,7 @@ DEF_OP(MemCpy) {
ARMEmitter::ForwardLabel DoneInternal {};
// Early exit if zero count.
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
ARMEmitter::ForwardLabel AbsPos {};
@@ -2174,11 +2174,11 @@ DEF_OP(MemCpy) {
ARMEmitter::BackwardLabel AgainInternal256 {};
sub(ARMEmitter::Size::i64Bit, TMP4, TMP2, TMP3);
tbz(TMP4, 63, &AbsPos);
(void)tbz(TMP4, 63, &AbsPos);
neg(ARMEmitter::Size::i64Bit, TMP4, TMP4);
Bind(&AbsPos);
(void)Bind(&AbsPos);
sub(ARMEmitter::Size::i64Bit, TMP4, TMP4, 32);
tbnz(TMP4, 63, &AgainInternal);
(void)tbnz(TMP4, 63, &AgainInternal);
if (Direction == -1) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
@@ -2190,30 +2190,30 @@ DEF_OP(MemCpy) {
// Do this in two parts, to fallback to the byte by byte loop if size < 32, and to the
// single copy loop if size < 64.
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbnz(TMP1, 63, &AgainInternal128Exit);
(void)tbnz(TMP1, 63, &AgainInternal128Exit);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbnz(TMP1, 63, &AgainInternal256Exit);
(void)tbnz(TMP1, 63, &AgainInternal256Exit);
Bind(&AgainInternal256);
(void)Bind(&AgainInternal256);
MemCpy(32, 32 * Direction);
MemCpy(32, 32 * Direction);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 64 / Size);
tbz(TMP1, 63, &AgainInternal256);
(void)tbz(TMP1, 63, &AgainInternal256);
Bind(&AgainInternal256Exit);
(void)Bind(&AgainInternal256Exit);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 64 / Size);
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbnz(TMP1, 63, &AgainInternal128Exit);
Bind(&AgainInternal128);
(void)tbnz(TMP1, 63, &AgainInternal128Exit);
(void)Bind(&AgainInternal128);
MemCpy(32, 32 * Direction);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
tbz(TMP1, 63, &AgainInternal128);
(void)tbz(TMP1, 63, &AgainInternal128);
Bind(&AgainInternal128Exit);
(void)Bind(&AgainInternal128Exit);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
@@ -2221,16 +2221,16 @@ DEF_OP(MemCpy) {
}
}
Bind(&AgainInternal);
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemCpyTSO(OpSize, SizeDirection);
} else {
MemCpy(OpSize, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
Bind(&DoneInternal);
(void)Bind(&DoneInternal);
// Needs to use temporaries just in case of overwrite
mov(TMP1, MemRegDest.X());
@@ -2288,11 +2288,11 @@ DEF_OP(MemCpy) {
for (int32_t Direction : {1, -1}) {
EmitMemcpy(Direction);
if (Direction == 1) {
b(&Done);
Bind(&BackwardImpl);
(void)b(&Done);
(void)Bind(&BackwardImpl);
}
}
Bind(&Done);
(void)Bind(&Done);
// Destination already set to the final pointer.
}
}
+34 -24
View File
@@ -1,79 +1,89 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
enum class RelocationTypes : uint8_t {
enum class RelocationTypes : uint32_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// 4 instruction constant generation
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// 8 byte literal (relative to binary base address)
RELOC_GUEST_RIP_LITERAL,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
// 4 instruction constant generation
// Aligned to struct RelocGuestRIP
RELOC_GUEST_RIP_MOVE,
};
struct RelocationTypeHeader final {
struct FEX_PACKED RelocationHeader final {
// Offset to the relocated host code data
uint64_t Offset {};
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
enum class NamedSymbol : uint32_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
RelocationTypeHeader Header {};
RelocationHeader Header {};
NamedSymbol Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
uint32_t Pad[8];
};
struct RelocNamedThunkMove final {
RelocationTypeHeader Header {};
RelocationHeader Header {};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
uint32_t RegisterIndex;
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
struct RelocGuestRIPMove final {
RelocationTypeHeader Header {};
struct RelocGuestRIP final {
RelocationHeader Header {};
// GPR index the constant is being moved to
// GPR index the constant is being moved to (for non-literal relocations)
uint8_t RegisterIndex;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
char Pad[3];
// The unrelocated RIP that is being moved
// The base RIP (to be moved by the register for non-literal relocations).
// In a serialized code cache, this is relative to the binary base address.
uint64_t GuestRIP;
uint32_t pad2[6] {};
};
union Relocation {
RelocationTypeHeader Header {};
RelocationHeader Header {};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIPMove GuestRIPMove;
RelocGuestRIP GuestRIP;
};
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl&, RelocNamedSymbolLiteral::NamedSymbol);
} // namespace FEXCore::CPU
@@ -39,6 +39,8 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We need one pointer per page of virtual memory
// At 64GB of virtual memory this will allocate 128MB of virtual memory space
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
@@ -49,14 +51,11 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = PagePointer + ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8;
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = PageMemory + CODE_SIZE;
FEXCore::Allocator::VirtualName("FEXMem_Lookup_L1", reinterpret_cast<void*>(L1Pointer), MAX_L1_SIZE);
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
if (DynamicL1Cache()) {
@@ -76,7 +75,7 @@ LookupCache::~LookupCache() {
// These will get freed when their memory allocators are deallocated.
}
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheWriteLockToken& lk) {
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer),
@@ -87,12 +86,12 @@ void LookupCache::ClearL2Cache(const FEXCore::LookupCacheWriteLockToken& lk) {
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
CachedCodePages.clear();
}
void LookupCache::ClearCache(const LookupCacheWriteLockToken& lk) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
ClearThreadLocalCaches(lk);
Shared->ClearCache(lk);
}
+95 -38
View File
@@ -3,9 +3,12 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/robin_set.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/memory_resource.h>
@@ -15,22 +18,41 @@
#include <mutex>
namespace FEXCore {
struct LookupCacheBaseLockToken {
protected:
// Protected constructor - only derived classes can construct
LookupCacheBaseLockToken() = default;
};
struct LookupCacheWriteLockToken {
struct LookupCacheWriteLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheWriteLockToken(std::mutex& Mutex)
LookupCacheWriteLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
: Lock {Mutex} {}
std::lock_guard<std::mutex> Lock;
std::lock_guard<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct LookupCacheReadLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheReadLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
: Lock {Mutex} {}
std::shared_lock<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct GuestToHostMap {
std::mutex WriteLock;
FEXCore::Utils::WritePriorityMutex::Mutex Lock {};
[[nodiscard]]
LookupCacheWriteLockToken AcquireWriteLock() {
return LookupCacheWriteLockToken {WriteLock};
return LookupCacheWriteLockToken {Lock};
}
[[nodiscard]]
LookupCacheReadLockToken AcquireReadLock() {
return LookupCacheReadLockToken {Lock};
}
struct BlockLinkTag {
@@ -59,42 +81,61 @@ struct GuestToHostMap {
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
struct BlockEntry {
uint64_t HostCode;
fextl::vector<uint64_t> CodePages;
};
fextl::robin_map<uint64_t, BlockEntry> BlockList;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap();
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode, const LookupCacheWriteLockToken&) {
const BlockEntry& AddBlockMapping(uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode, const LookupCacheWriteLockToken&) {
// This may replace an existing mapping
// NOTE: Generally no previous entry should exist, however there is one exception:
// If the backend updates the active thread's CodeBuffer, the new associated LookupCache
// may already contain the block address. Since is comparatively rare, we'll just leak
// one of the two blocks in this case.
BlockList[Address] = (uintptr_t)HostCode;
return BlockList.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, CodePages}).first->second;
}
std::optional<uintptr_t> FindBlock(uint64_t Address, const LookupCacheWriteLockToken&) {
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
auto HostCode = BlockList.find(Address);
if (HostCode == BlockList.end()) {
return std::nullopt;
return nullptr;
}
return HostCode->second;
return &HostCode->second;
}
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken&) {
bool Erase(uint64_t Address, const LookupCacheWriteLockToken&) {
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
it->second(it->first.HostLink);
}
// Remove from BlockList
return BlockList.erase(Address) != 0;
}
void InvalidateRange(uint64_t Start, uint64_t Length) {
auto lk = AcquireWriteLock();
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
Erase(Entry, lk);
}
}
CodePages.erase(lower, upper);
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
const FEXCore::Context::BlockDelinkerFunc& delinker, const LookupCacheWriteLockToken&) {
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
@@ -144,7 +185,7 @@ public:
{
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheReadLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
auto lk = Shared->AcquireReadLock();
LockTime.reset();
if (!DisableL2Cache()) {
@@ -170,10 +211,10 @@ public:
if (!HostPtr) {
// Try L3
auto HostCode = Shared->FindBlock(Address, lk);
if (HostCode) {
CacheBlockMapping(Address, HostCode.value(), lk);
HostPtr = HostCode.value();
auto Entry = Shared->FindBlock(Address, lk);
if (Entry) {
CacheBlockMapping(Address, *Entry, false, lk);
HostPtr = Entry->HostCode;
}
}
}
@@ -244,32 +285,25 @@ public:
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, void* HostCode) {
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
Shared->AddBlockMapping(Address, HostCode, lk);
const auto& Entry = Shared->AddBlockMapping(Address, CodePages, HostCode, lk);
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
CacheBlockMapping(Address, Entry, true, lk);
}
// NOTE: It's the caller's responsibility to call Erase() for all other
// GuestToHostMaps that share the same LookupCache. Otherwise, the
// L1/L2 caches will contain stale references to deallocated memory.
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken& lk) {
bool ErasedAny = Shared->Erase(Frame, Address, lk);
// Invalidates L1/L2 for a given guest block
void InvalidateCache(uint64_t Address, const LookupCacheWriteLockToken& lk) {
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
ErasedAny = true;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
@@ -285,7 +319,7 @@ public:
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return ErasedAny;
return;
}
// Page exists, just set the offset to zero
@@ -293,7 +327,23 @@ public:
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
}
return true;
}
// Invalidates all L1/L2 entries for all guest block that intersect the given range
bool InvalidateCacheRange(uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireWriteLock();
auto lower = CachedCodePages.lower_bound(Start >> 12);
auto upper = CachedCodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
InvalidateCache(Entry, lk);
}
}
bool ret = upper != lower;
CachedCodePages.erase(lower, upper);
return ret;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
@@ -302,7 +352,7 @@ public:
}
void ClearCache(const LookupCacheWriteLockToken&);
void ClearL2Cache(const LookupCacheWriteLockToken&);
void ClearL2Cache(const LookupCacheBaseLockToken&);
void ClearThreadLocalCaches(const LookupCacheWriteLockToken&);
uintptr_t GetL1Pointer() const {
@@ -330,13 +380,17 @@ public:
}
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode, const LookupCacheWriteLockToken& lk) {
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheBaseLockToken& lk) {
for (const auto& CodePage : Entry.CodePages) {
CachedCodePages[CodePage >> 12].insert(Address);
}
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = HostCode;
L1Entry.HostCode = Entry.HostCode;
if (!DisableL2Cache()) {
if (!DisableL2Cache() && !L1Only) {
// Do ful map
auto FullAddress = Address;
Address = Address & (VirtualMemSize - 1);
@@ -353,7 +407,7 @@ private:
if (!NewPageBacking) {
// Couldn't allocate, clear L2 and retry
ClearL2Cache(lk);
CacheBlockMapping(Address, HostCode, lk);
CacheBlockMapping(Address, Entry, false, lk);
return;
}
Pointers[Address] = NewPageBacking;
@@ -365,7 +419,7 @@ private:
// This silently replaces existing mappings
BlockPointers[PageOffset].GuestCode = FullAddress;
BlockPointers[PageOffset].HostCode = HostCode;
BlockPointers[PageOffset].HostCode = Entry.HostCode;
}
}
@@ -383,6 +437,9 @@ private:
return PageMemory + NewBase;
}
// Maps from a page index to all blocks in the page that have at some point been fetched into L1/L2
fextl::map<uint64_t, fextl::robin_set<uint64_t>> CachedCodePages;
uintptr_t PagePointer;
uintptr_t PageMemory;
uintptr_t L1Pointer;
@@ -514,18 +514,17 @@ void OpDispatchBuilder::CALLOp(OpcodeArgs) {
BlockSetRIP = true;
// Call instruction only uses up to 32-bit signed displacement
int64_t TargetOffset = Op->Src[0].Literal();
const int64_t TargetOffset = Op->Src[0].Literal();
auto ConstantPC = GetRelocatedPC(Op);
const auto ConstantPC = GetRelocatedPC(Op);
// Push the return address.
Push(GPRSize, ConstantPC);
const uint64_t NextRIP = Op->PC + Op->InstSize;
uint64_t TargetRIP = NextRIP + TargetOffset;
if (NextRIP != TargetRIP) {
if (TargetOffset != 0) {
// Store the RIP
const uint64_t NextRIP = Op->PC + Op->InstSize;
ExitRelocatedPC(Op, TargetOffset, BranchHint::Call, ConstantPC, [&]() {
auto CallReturnJumpTarget = JumpTargets.find(NextRIP);
if (CallReturnJumpTarget != JumpTargets.end() && CallReturnJumpTarget->second.IsEntryPoint) {
@@ -2750,7 +2749,8 @@ void OpDispatchBuilder::NOTOp(OpcodeArgs) {
if (DestIsLockedMem(Op)) {
HandledLock = true;
Ref DestMem = MakeSegmentAddress(Op, Op->Dest);
_AtomicXor(Size, MaskConst, DestMem);
// Result unused
_AtomicFetchXor(Size, MaskConst, DestMem);
} else if (!Op->Dest.IsGPR()) {
// GPR version plays fast and loose with sizes, be safe for memory tho.
Ref Src = LoadSourceGPR(Op, Op->Dest, Op->Flags);
@@ -3199,7 +3199,7 @@ void OpDispatchBuilder::DECOp(OpcodeArgs) {
void OpDispatchBuilder::STOSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("STOSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3244,7 +3244,7 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("MOVSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3297,45 +3297,57 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
_StoreMem(RegClass::GPR, Size, Src, RDI, Invalid(), OpSize::i8Bit, MemOffsetType::SXTX, 1);
}
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
RSI = Add(OpSize::i64Bit, RSI, PtrDir);
RDI = Add(OpSize::i64Bit, RDI, PtrDir);
RSI = OffsetByDir(RSI, IR::OpSizeToSize(Size));
RDI = OffsetByDir(RDI, IR::OpSizeToSize(Size));
StoreGPRRegister(X86State::REG_RSI, RSI);
StoreGPRRegister(X86State::REG_RDI, RDI);
}
}
IR::OpSize OpDispatchBuilder::GetStringOpSize(X86Tables::DecodedOp Op) const {
LOGMAN_THROW_A_FMT(Is64BitMode || !(Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE), "Invalid modifier on 32bit address");
return !Is64BitMode || (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) ? OpSize::i32Bit : OpSize::i64Bit;
}
void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("CMPSOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
bool Repeat = Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX);
if (!Repeat) {
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src2, Src1);
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
Dest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
Dest_RSI = OffsetByDir(Src_RSI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
} else {
// Calculate flags early.
CalculateDeferredFlags();
@@ -3351,7 +3363,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
SetCurrentCodeBlock(BeforeLoop);
StartNewBlock();
ForeachDirection([this, Op, Size, REPE](int32_t PtrDir) {
ForeachDirection([this, Op, Size, AddrSize, REPE](int32_t PtrDir) {
IRPair<IROp_CondJump> InnerJump;
auto JumpIntoLoop = Jump();
@@ -3363,10 +3375,11 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Working loop
{
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPR(Size, Dest_RSI, Size);
@@ -3383,13 +3396,21 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
Dest_RDI = Add(AddrSize, Src_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
Dest_RSI = Add(AddrSize, Src_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
// If TailCounter != 0, compare sources.
// If TailCounter == 0, set ZF iff that would break.
@@ -3428,7 +3449,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("LODSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3510,31 +3531,37 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
}
void OpDispatchBuilder::SCASOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("SCASOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX)) != 0;
if (!Repeat) {
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src1, Src2);
// Offset the pointer
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
StoreGPRRegister(X86State::REG_RDI, OffsetByDir(TailDest_RDI, IR::OpSizeToSize(Size)));
Ref TailDest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
} else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
ForeachDirection([this, Op, Size](int32_t Dir) {
ForeachDirection([this, Op, Size, AddrSize](int32_t Dir) {
bool REPE = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX;
auto JumpStart = Jump();
@@ -3558,7 +3585,8 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Working loop
{
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
@@ -3569,7 +3597,7 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
CalculateDeferredFlags();
Ref TailCounter = LoadGPRRegister(X86State::REG_RCX);
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
Ref Src_RDI_Tail = LoadGPRRegister(X86State::REG_RDI, AddrSize);
// Decrement counter
TailCounter = Sub(OpSize::i64Bit, TailCounter, 1);
@@ -3577,9 +3605,13 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
TailDest_RDI = Add(OpSize::i64Bit, TailDest_RDI, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
Ref TailDest_RDI = Add(AddrSize, Src_RDI_Tail, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
CalculateDeferredFlags();
InternalCondJump = CondJumpNZCV(REPE ? CondClass::EQ : CondClass::NEQ);
@@ -1655,6 +1655,9 @@ private:
return IR::SizeToOpSize(GetSrcSize(Op));
}
[[nodiscard]]
IR::OpSize GetStringOpSize(X86Tables::DecodedOp Op) const;
// Set flag tracking to prepare for an operation that directly writes NZCV.
void HandleNZCVWrite() {
CachedNZCV = nullptr;
@@ -552,7 +552,7 @@ void OpDispatchBuilder::AVXInsertScalarRound(OpcodeArgs) {
const uint64_t Mode = Op->Src[2].Literal();
const auto DstSize = GetGuestVectorLength();
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Dest, Op->Src[0], Mode, true);
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Src[0], Op->Src[1], Mode, true);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -17,7 +17,6 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/FPState.h>
#include <cmath>
#include <stddef.h>
#include <stdint.h>
@@ -61,15 +60,13 @@ void OpDispatchBuilder::SetX87Top(Ref Value) {
// Float LoaD operation with memory operand
void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
Ref ConvertedData = Data;
// Convert to 80bit float
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
ConvertedData = _F80CVTTo(Data, ReadWidth);
ConvertedData = _F80CVTTo(Data, Width);
}
_PushStack(ConvertedData, Data, ReadWidth, true);
_PushStack(ConvertedData, Data, Width);
}
// Float LoaD operation with memory operand
@@ -81,7 +78,7 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Data, OpSize::i128Bit, true);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
@@ -93,7 +90,7 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::i128Bit, true);
_PushStack(Data, Data, OpSize::f80Bit);
}
void OpDispatchBuilder::FILD(OpcodeArgs) {
@@ -124,15 +121,16 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VLoadTwoGPRs(shifted, upper);
_PushStack(ConvertedData, Data, ReadWidth, false);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Width == OpSize::i32Bit || Width == OpSize::i64Bit || Width == OpSize::f80Bit, "Invalid store width for FST");
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::f80Bit;
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale, /*Float=*/true);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
_PopStackDestroy();
@@ -878,8 +876,8 @@ void OpDispatchBuilder::X87FXTRACT(OpcodeArgs) {
_PopStackDestroy();
auto Exp = _F80XTRACT_EXP(Top);
auto Sig = _F80XTRACT_SIG(Top);
_PushStack(Exp, Exp, OpSize::f80Bit, true);
_PushStack(Sig, Sig, OpSize::f80Bit, true);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -59,7 +59,6 @@ void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
// F64 ops
// Float load op with memory operand
void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
// Convert to 64bit float
Ref ConvertedData = Data;
@@ -68,7 +67,7 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
} else if (Width == OpSize::f80Bit) {
ConvertedData = _F80CVT(OpSize::i64Bit, Data);
}
_PushStack(ConvertedData, Data, ReadWidth, true);
_PushStack(ConvertedData, Data, Width);
}
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
@@ -76,7 +75,7 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Data, OpSize::i64Bit, true);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
@@ -88,7 +87,7 @@ void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
void OpDispatchBuilder::FLDF64_Const(OpcodeArgs, uint64_t Num) {
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(Num));
_PushStack(Data, Data, OpSize::i64Bit, true);
_PushStack(Data, Data, OpSize::i64Bit);
}
void OpDispatchBuilder::FILDF64(OpcodeArgs) {
@@ -100,7 +99,7 @@ void OpDispatchBuilder::FILDF64(OpcodeArgs) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
auto ConvertedData = _Float_FromGPR_S(OpSize::i64Bit, ReadWidth == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, Data);
_PushStack(ConvertedData, Data, ReadWidth, false);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
@@ -397,7 +396,7 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
_PopStackDestroy();
_PushStack(Exp, Exp, OpSize::i64Bit, true);
_PushStack(Sig, Sig, OpSize::i64Bit, true);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -1,89 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
desc: Guest-side assembly helpers used by the backends
$end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#include <cstring>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
X86GeneratedCode::X86GeneratedCode() {
#ifdef _WIN32
// No need to allocate anything in this config.
#else
// Allocate a page for our emulated guest
CodePtr = AllocateGuestCodeSpace(CODE_SIZE);
constexpr std::array<uint8_t, 2> SignalReturnCode = {
0x0F, 0x3E, // CALLBACKRET FEX Instruction
};
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
memcpy(reinterpret_cast<void*>(CallbackReturn), SignalReturnCode.data(), SignalReturnCode.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
#endif
}
X86GeneratedCode::~X86GeneratedCode() {
#ifndef _WIN32
FEXCore::Allocator::VirtualFree(CodePtr, CODE_SIZE);
#endif
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
#ifndef _WIN32
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
auto Result = FEXCore::Allocator::VirtualAlloc(Size);
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(Result), Size);
return Result;
}
// First 64bit page
constexpr uintptr_t LOCATION_MAX = 0x1'0000'0000;
// 32bit mode
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void* Ptr = ::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED && reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
::munmap(Ptr, Size);
continue;
}
if (Ptr != MAP_FAILED) {
return Ptr;
}
}
// Can't do anything about this
// Here's hoping the application doesn't use signals
return MAP_FAILED;
#else
return nullptr;
#endif
}
} // namespace FEXCore
@@ -1,25 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
$end_info$
*/
#pragma once
#include <stddef.h>
#include <stdint.h>
namespace FEXCore {
class X86GeneratedCode final {
public:
X86GeneratedCode();
~X86GeneratedCode();
uint64_t CallbackReturn {};
private:
void* CodePtr {};
void* AllocateGuestCodeSpace(size_t Size);
};
} // namespace FEXCore
+1 -18
View File
@@ -60,24 +60,7 @@ struct NodeID final {
Value = 0;
}
[[nodiscard]] friend constexpr bool operator==(NodeID, NodeID) noexcept = default;
[[nodiscard]]
friend constexpr bool operator<(NodeID lhs, NodeID rhs) noexcept {
return lhs.Value < rhs.Value;
}
[[nodiscard]]
friend constexpr bool operator>(NodeID lhs, NodeID rhs) noexcept {
return operator<(rhs, lhs);
}
[[nodiscard]]
friend constexpr bool operator<=(NodeID lhs, NodeID rhs) noexcept {
return !operator>(lhs, rhs);
}
[[nodiscard]]
friend constexpr bool operator>=(NodeID lhs, NodeID rhs) noexcept {
return !operator<(lhs, rhs);
}
[[nodiscard]] constexpr auto operator<=>(const NodeID&) const noexcept = default;
friend std::ostream& operator<<(std::ostream& out, NodeID ID) {
out << ID.Value;
+5 -20
View File
@@ -804,16 +804,6 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AtomicXor OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer xor",
"IR layout must match Fetch-variant, otherwise DCE IR optimization breaks!"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i8Bit || Size == FEXCore::IR::OpSize::i16Bit || Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AtomicSwap OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer swap"
@@ -2818,17 +2808,13 @@
"X87": true,
"HasSideEffects": true
},
"PushStack FPR:$X80Src, SSA:$OriginalValue, OpSize:$LoadSize, i1:$Float": {
"PushStack FPR:$X80Src, FPR:$OriginalValue, OpSize:$LoadSize": {
"Desc": [
"Pushes the provided X80Src source on to the x87 stack.",
"Tracks OriginalValue as the original value of X80Src.",
"Tracks OriginalValue as the original value of X80Src. OriginalValue can be Invalid() in which case no tracking is done.",
"Opsize is 128bit for F80 values, 64-bit for low precision.",
"LoadSize the original load size, i.e. of size of OriginalValue.",
"Float: 80-bit, 64-bit, 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"EmitValidation": [
"WalkFindRegClass($OriginalValue) == RegClass::FPR || WalkFindRegClass($OriginalValue) == RegClass::GPR"
"Float: 80-bit, 64-bit, 32-bit"
],
"HasSideEffects": true,
"X87": true
@@ -2840,13 +2826,12 @@
"HasSideEffects": true,
"X87": true
},
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale, i1:$Float": {
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale": {
"Desc": [
"Takes the top value off the x87 stack and stores it to memory.",
"SourceSize is 128bit for F80 values, 64-bit for low precision.",
"StoreSize is the store size for conversion:",
"Float: 80-bit, 64-bit, or 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
"Float: 80-bit, 64-bit, or 32-bit"
],
"HasSideEffects": true,
"X87": true
+1 -2
View File
@@ -353,8 +353,7 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
AddIndent();
*out << "(%" << IR->GetID(BlockNode) << ") "
<< "CodeBlock ";
*out << "(%" << IR->GetID(BlockNode) << ") " << "CodeBlock ";
*out << "%" << BlockIROp->Begin.ID() << ", ";
*out << "%" << BlockIROp->Last.ID() << std::endl;
@@ -6,7 +6,6 @@
#include "Interface/IR/PassManager.h"
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/Profiler.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/Core/HostFeatures.h"
#include "Interface/Core/Addressing.h"
@@ -66,7 +65,7 @@ public:
int8_t TopOffset = 0;
FixedSizeStack()
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T()}) {}
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T::Invalid}) {}
void push(const T& Value) {
rotate();
@@ -85,7 +84,7 @@ public:
}
void pop() {
buffer.front() = {StackSlot::INVALID, T()};
buffer.front() = {StackSlot::INVALID, T::Invalid};
rotate(false);
}
@@ -103,7 +102,7 @@ public:
void clear() {
for (auto& Elem : buffer) {
Elem = {StackSlot::UNUSED, T()};
Elem = {StackSlot::UNUSED, T::Invalid};
}
TopOffset = 0;
}
@@ -171,13 +170,8 @@ private:
// Helpers
Ref RotateRight8(uint32_t V, Ref Amount);
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
IREmit->_StoreMemFPR(OpSize::i64Bit, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
auto Upper = IREmit->_VExtractToGPR(OpSize::i128Bit, OpSize::i64Bit, StackNode, 1);
@@ -192,7 +186,24 @@ private:
IREmit->_StoreMemGPR(OpSize::i16Bit, Upper, A.Base, A.Index, OpSize::i64Bit, MemOffsetType::SXTX, A.IndexScale);
}
void Store80BitToMem(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else {
F80SplitStore_Helper(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
}
void StoreStackMem_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(!ReducedPrecisionMode, "Full precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -209,17 +220,7 @@ private:
}
case OpSize::f80Bit: {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else { // 80bit requires split-store
F80SplitStore_Helper(Op, StackNode);
}
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -229,6 +230,8 @@ private:
// Performs a store to memory from a value the stack passed in as StackNode.
// This is the version dealing with the reduced precision case.
void StoreStackMem_Reduced_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(ReducedPrecisionMode, "Reduced precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -245,10 +248,9 @@ private:
break;
}
// 80bit requires split-store
case OpSize::f80Bit: {
StackNode = IREmit->_F80CVTTo(StackNode, OpSize::i64Bit);
F80SplitStore_Helper(Op, StackNode);
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -290,23 +292,24 @@ private:
void Reset();
struct StackMemberInfo {
StackMemberInfo() {}
StackMemberInfo() = delete;
StackMemberInfo(Ref Data)
: StackDataNode(Data) {}
StackMemberInfo(Ref Data, Ref Source, OpSize Size, bool Float)
StackMemberInfo(Ref Data, Ref Source, OpSize Size)
: StackDataNode(Data)
, Source({Size, Source})
, InterpretAsFloat(Float) {}
, Source({Size, Source}) {}
Ref StackDataNode {}; // Reference to the data in the Stack.
// This is the source data node in the stack format, possibly converted to 64/80 bits.
struct StackMemberData final {
OpSize Size;
Ref Node;
};
static const StackMemberInfo Invalid;
// Tuple is only valid if we have information about the Source of the Stack Data Node.
// In it's valid then OpSize is the original source size and Ref is the original source node.
std::optional<StackMemberData> Source {};
bool InterpretAsFloat {false}; // True if this is a floating point value, false if integer
};
// StackData, TopCache need to be always properly set to ensure
@@ -359,6 +362,8 @@ private:
IRListView* IR = nullptr;
};
inline const X87StackOptimization::StackMemberInfo X87StackOptimization::StackMemberInfo::Invalid {nullptr};
inline void X87StackOptimization::InvalidateCaches() {
InvalidateCachedRegs();
ConstantPool.fill(nullptr);
@@ -728,6 +733,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// The optimization should run per-block
Reset();
IREmit->SetCurrentCodeBlock(BlockNode);
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (!LoweredX87(IROp->Op)) {
continue;
@@ -927,8 +933,13 @@ void X87StackOptimization::Run(IREmitter* Emit) {
StoreStackValueAtOffset_Slow(SourceNode);
} else {
auto* SourceNode = CurrentIR.GetNode(Op->X80Src);
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize, Op->Float});
if (Op->OriginalValue.IsInvalid()) {
// No original value to track - just push the converted data
StackData.push(StackMemberInfo {SourceNode});
} else {
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize});
}
}
break;
}
@@ -993,9 +1004,16 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// str w2, [x1]
// or similar. As long as the source size and dest size are one and the same.
// This will avoid any conversions between source and stack element size and conversion back.
if (!SlowPath && Value->Source && Value->Source->Size == Op->StoreSize && Value->InterpretAsFloat) {
const auto ClassType = Value->InterpretAsFloat ? RegClass::FPR : RegClass::GPR;
IREmit->_StoreMem(ClassType, Op->StoreSize, Value->Source->Node, AddrNode, Offset, Align, OffsetType, OffsetScale);
OpSize StoreSize = Op->StoreSize;
LOGMAN_THROW_A_FMT(Op->StoreSize == OpSize::i32Bit || Op->StoreSize == OpSize::i64Bit || Op->StoreSize == OpSize::f80Bit,
"Invalid store size in x87 store stack mem");
if (!SlowPath && Value->Source && Value->Source->Size == StoreSize) {
Ref SourceValue = Value->Source->Node;
if (Op->StoreSize == OpSize::f80Bit) {
Store80BitToMem(Op, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
} else {
IREmit->_StoreMemFPR(StoreSize, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
break;
}
@@ -1035,11 +1053,26 @@ void X87StackOptimization::Run(IREmitter* Emit) {
case OP_F80STACKXCHANGE: {
const auto* Op = IROp->C<IROp_F80StackXchange>();
auto Offset = Op->SrcStack;
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
if (Offset == 0) {
// No-op
break;
}
const auto [ValidTop, StackMemberTop] = StackData.top(0);
const auto [ValidOffset, StackMemberOffset] = StackData.top(Offset);
if (ValidTop != StackSlot::VALID || ValidOffset != StackSlot::VALID) {
// Slow path: do actual memory operations
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
} else {
// Fast path: swap complete StackMemberInfo preserving Source metadata
StackData.setTop(StackMemberOffset, 0);
StackData.setTop(StackMemberTop, Offset);
}
break;
}
+1 -4
View File
@@ -140,10 +140,7 @@ void ClearHooks() {
FEXCore::Allocator::mmap = ::mmap;
FEXCore::Allocator::munmap = ::munmap;
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Alloc64.release();
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Alloc64));
}
#pragma GCC diagnostic pop
@@ -207,7 +207,7 @@ OSAllocator_64Bit::LiveVMARegion* OSAllocator_64Bit::FindLiveRegionForAddress(ui
uintptr_t RegionBegin = (*it)->SlabInfo->Base;
uintptr_t RegionEnd = RegionBegin + (*it)->SlabInfo->RegionSize;
if (Addr >= RegionBegin && Addr < RegionEnd) {
if (Addr >= RegionBegin && AddrEnd < RegionEnd) {
LiveRegion = *it;
// Leave our loop
break;
@@ -405,14 +405,18 @@ again:
// Mark the pages as used
uintptr_t RegionBegin = LiveRegion->SlabInfo->Base;
uintptr_t MappedBegin = (AllocatedOffset - RegionBegin) >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t PagesSet {};
for (size_t i = 0; i < NumberOfPages; ++i) {
LiveRegion->UsedPages.Set(MappedBegin + i);
PagesSet += LiveRegion->UsedPages.TestAndSet(MappedBegin + i) == false;
}
// Change our last allocation region
LiveRegion->LastPageAllocation = MappedBegin + NumberOfPages;
LiveRegion->FreeSpace -= length;
LiveRegion->FreeSpace -= PagesSet * FEXCore::Utils::FEX_PAGE_SIZE;
LOGMAN_THROW_A_FMT(LiveRegion->FreeSpace <= LiveRegion->SlabInfo->RegionSize,
"Corrupt LiveRegion free space! 0x{:x} > 0x{:x}. After allocating 0x{:x} (0x{:x} overlapped)", LiveRegion->FreeSpace,
LiveRegion->SlabInfo->RegionSize, length, PagesSet);
}
if (!AllocatedOffset) {
+20 -6
View File
@@ -27,6 +27,11 @@ struct FlexBitSet final {
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
return Value;
}
bool TestAndSet(size_t Element) {
bool Value = Get(Element);
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
return Value;
}
void Set(size_t Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
@@ -70,12 +75,17 @@ struct FlexBitSet final {
template<bool WantUnset>
BitsetScanResults BackwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t MinimumElement) {
bool FoundHole {};
for (size_t CurrentPage = BeginningElement; CurrentPage >= (MinimumElement + ElementCount);) {
// Final element to iterate to.
const size_t FinalElement = MinimumElement + ElementCount - 1;
for (size_t CurrentPage = BeginningElement; CurrentPage >= FinalElement;) {
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT(Remaining <= CurrentPage, "Scanning less than available range");
LOGMAN_THROW_A_FMT(CurrentPage <= BeginningElement && CurrentPage >= FinalElement, "BackwardScanForRange: Scanning less than "
"available range");
while (Remaining) {
if (this->Get(CurrentPage - Remaining) == WantUnset) {
if (this->Get(CurrentPage - Remaining + 1) == WantUnset) {
// Has an intersecting range
break;
}
@@ -92,7 +102,7 @@ struct FlexBitSet final {
CurrentPage -= Remaining;
} else {
// We have a slab range
return BitsetScanResults {CurrentPage - ElementCount, FoundHole};
return BitsetScanResults {CurrentPage - ElementCount + 1, FoundHole};
}
}
@@ -108,11 +118,15 @@ struct FlexBitSet final {
BitsetScanResults ForwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t ElementsInSet) {
bool FoundHole {};
for (size_t CurrentElement = BeginningElement; CurrentElement < (ElementsInSet - ElementCount);) {
// Final element to iterate to.
const size_t FinalElement = ElementsInSet - ElementCount + 1;
for (size_t CurrentElement = BeginningElement; CurrentElement <= FinalElement;) {
// If we have enough free space, check if we have enough free pages that are contiguous
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT((CurrentElement + Remaining - 1) < ElementsInSet, "Scanning less than available range");
LOGMAN_THROW_A_FMT(CurrentElement >= BeginningElement && CurrentElement <= FinalElement, "ForwardScanForRange: Scanning less than "
"available range");
while (Remaining) {
if (this->Get(CurrentElement + Remaining - 1) == WantUnset) {
@@ -53,4 +53,12 @@ public:
namespace Alloc::OSAllocator {
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocator();
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocatorWithRegions(fextl::vector<FEXCore::Allocator::MemoryRegion>& Regions);
static inline void ReleaseAllocatorWorkaround(fextl::unique_ptr<Alloc::HostAllocator> Allocator) {
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// The allocator is also intrusively allocated, so the unique_ptr tries to double free the HostAllocator object.
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Allocator.release();
}
} // namespace Alloc::OSAllocator
+147
View File
@@ -0,0 +1,147 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
namespace FEXCore::UncheckedLongJump {
#if defined(_M_ARM_64)
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
__asm volatile(R"(
// x0 contains the jumpbuffer
stp x19, x20, [x0, #( 0 * 8)];
stp x21, x22, [x0, #( 2 * 8)];
stp x23, x24, [x0, #( 4 * 8)];
stp x25, x26, [x0, #( 6 * 8)];
stp x27, x28, [x0, #( 8 * 8)];
stp x29, x30, [x0, #(10 * 8)];
// FPRs
stp d8, d9, [x0, #(12 * 8)];
stp d10, d11, [x0, #(14 * 8)];
stp d12, d13, [x0, #(16 * 8)];
stp d14, d15, [x0, #(18 * 8)];
// Move SP in to a temporary to store.
mov x1, sp;
str x1, [x0, #(20 * 8)];
// Return zero to signify this is the SetJump.
mov x0, #0;
ret;
)" ::
: "memory");
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
// x0 contains the jumpbuffer
ldp x19, x20, [x0, #( 0 * 8)];
ldp x21, x22, [x0, #( 2 * 8)];
ldp x23, x24, [x0, #( 4 * 8)];
ldp x25, x26, [x0, #( 6 * 8)];
ldp x27, x28, [x0, #( 8 * 8)];
ldp x29, x30, [x0, #(10 * 8)];
// FPRs
ldp d8, d9, [x0, #(12 * 8)];
ldp d10, d11, [x0, #(14 * 8)];
ldp d12, d13, [x0, #(16 * 8)];
ldp d14, d15, [x0, #(18 * 8)];
// Load SP in to temporary then move
ldr x0, [x0, #(20 * 8)];
mov sp, x0;
// Move value in to result register
mov x0, x1;
ret;
)" ::
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
// First 12 values are registers [x19,x30].
memcpy(&GPRs[19], &Buffer.Registers[0], sizeof(uint64_t) * 12);
// Next 8 values are [D8,D15]
// Retain upper 64-bits of the register, only modifying lower 64-bits.
for (size_t i = 0; i < 8; ++i) {
memcpy(&FPRs[8 + i], &Buffer.Registers[12 + i], sizeof(uint64_t));
}
// Last value is stack pointer
memcpy(&GPRs[31], &Buffer.Registers[20], sizeof(uint64_t));
// Load the expected value in to X0
GPRs[0] = Value;
// Load the PC with the current LR.
*PC = GPRs[30];
}
#else
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
__asm volatile(R"(
.intel_syntax noprefix;
// rdi contains the jumpbuffer
mov [rdi + (0 * 8)], rbx;
mov [rdi + (1 * 8)], rsp;
mov [rdi + (2 * 8)], rbp;
mov [rdi + (3 * 8)], r12;
mov [rdi + (4 * 8)], r13;
mov [rdi + (5 * 8)], r14;
mov [rdi + (6 * 8)], r15;
// Return address is on the stack, load it and store
mov rsi, [rsp];
mov [rdi + (7 * 8)], rsi;
// Return zero to signify this is the SetJump.
mov rax, 0;
ret;
.att_syntax prefix;
)" ::
: "memory");
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
.intel_syntax noprefix;
// rdi contains the jumpbuffer
mov rbx, [rdi + (0 * 8)];
mov rsp, [rdi + (1 * 8)];
mov rbp, [rdi + (2 * 8)];
mov r12, [rdi + (3 * 8)];
mov r13, [rdi + (4 * 8)];
mov r14, [rdi + (5 * 8)];
mov r15, [rdi + (6 * 8)];
// Move value in to result register
mov rax, rsi;
// Pop the dead return address off the stack
pop rsi;
// Load the original return address from the jumpbuffer
mov rsi, [rdi + (7 * 8)];
// Return using a jump
jmp rsi;
.att_syntax prefix;
)" ::
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
LOGMAN_MSG_A_FMT("This is unimplemented on x86-64");
}
#endif
} // namespace FEXCore::UncheckedLongJump
+45 -23
View File
@@ -1,9 +1,14 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <chrono>
#include <mutex>
#include <type_traits>
#include <FEXCore/fextl/functional.h>
#include <FEXCore/Utils/EnumUtils.h>
namespace FEXCore::Utils::SpinWaitLock {
/**
* @brief This provides routines to implement implement an "efficient spin-loop" using ARM's WFE and exclusive monitor interfaces.
@@ -123,29 +128,20 @@ static inline uint64_t WFELoadAtomic(uint64_t* Futex) {
return Result;
}
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
return;
}
do {
while (!Pred {}(Result, ComparisonValue)) {
Result = LoadExclusive(Futex);
if (Result == ExpectedValue) {
if (Pred {}(Result, ComparisonValue)) {
return;
}
Result = WFELoadAtomic(Futex);
} while (Result != ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
Result = WFELoadAtomic(Futex);
}
}
template<typename T, typename TT>
static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanoseconds& Timeout) {
@@ -184,20 +180,36 @@ template bool Wait<uint16_t>(uint16_t*, uint16_t, const std::chrono::nanoseconds
template bool Wait<uint32_t>(uint32_t*, uint32_t, const std::chrono::nanoseconds&);
template bool Wait<uint64_t>(uint64_t*, uint64_t, const std::chrono::nanoseconds&);
#else
template<typename T, typename TT>
static inline void Wait(T* Futex, TT ExpectedValue) {
template<typename T>
static inline T OneShotWFEBitComparison(T* Futex, T Mask, T Comp) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
return;
if ((Result & Mask) == Comp) {
return Result;
}
do {
Result = LoadExclusive(Futex);
if ((Result & Mask) == Comp) {
return Result;
}
// Waits for write and returns result.
Result = WFELoadAtomic(Futex);
return Result;
}
#else
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
while (!Pred {}(Result, ComparisonValue)) {
Result = AtomicFutex.load();
} while (Result != ExpectedValue);
}
}
template<typename T, typename TT>
@@ -228,6 +240,16 @@ static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanosecon
}
#endif
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
WaitPred<std::equal_to<>, T>(Futex, ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
template<typename T>
static inline void lock(T* Futex) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
+381
View File
@@ -0,0 +1,381 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <cstdint>
#if !defined(_WIN32)
#include <linux/futex.h> /* Definition of FUTEX_* constants */
#include <sys/syscall.h> /* Definition of SYS_* constants */
#include <unistd.h>
#else
#include <synchapi.h>
#endif
#include <FEXCore/Utils/LogManager.h>
#include "Utils/SpinWaitLock.h"
namespace FEXCore::Utils::WritePriorityMutex {
// A custom mutex that prioritizes exclusive locks.
// In highly contested scenarios, this can help minimize overall contention time.
//
// Features:
// - Up to 32767 pending exclusive locks ("writers")
// - Up to 32767 pending shared_locks ("readers")
// - Low-overhead waiting via WFE with a fallback to futex on timeout
// - Direct writer->reader hand-off and vice-versa to further reduce overhead
//
// Trade-offs:
// - No guaranteed order of wake-ups besides prioritizing writers
// - No support for recursive locking
// - We can't use FUTEX_LOCK_PI to enable priority inheritance
class Mutex final {
public:
Mutex() = default;
// Move-only type
Mutex(const Mutex&) = delete;
Mutex& operator=(const Mutex&) = delete;
Mutex(Mutex&& rhs) = delete;
Mutex& operator=(Mutex&&) = delete;
void lock() {
// Try a non-blocking lock first.
if (try_lock()) {
return;
}
// Try a quick WFE write-lock.
if (Attempt_WFE_WriteLock()) {
return;
}
// Still couldn't get it. Start waiting.
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected {};
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
// Increment the number of write waiters.
Desired = Expected + WRITE_WAITER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & WRITE_WAITER_COUNT_MASK) != 0, "Overflow in write-waiters!");
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Expected = AtomicFutex.fetch_add(WRITE_WAITER_INCREMENT);
Desired = Expected + WRITE_WAITER_INCREMENT;
#endif
// Thread added to waiter list.
Expected = Desired;
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & READ_OWNER_COUNT_MASK) == 0) {
// If not write-owned, and no read-owners, try to acquire.
LOGMAN_THROW_A_FMT((Expected & WRITE_WAITER_COUNT_MASK) != 0, "Underflow in write-waiters!");
// Add write-owned bit.
Desired = Expected | WRITE_OWNED_BIT;
// Remove ourselves from the wait list.
Desired -= WRITE_WAITER_INCREMENT;
Sleep = false;
} else {
// Already write-owned or read-locked. Go to sleep.
Desired = Expected;
Sleep = true;
break;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Somehow acquired a write-lock without it being set!");
return;
}
FutexWaitForWriteAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void lock_shared() {
// Try an uncontended lock first.
if (try_lock_shared()) {
return;
}
// Try a quick WFE read-lock.
if (Attempt_WFE_ReadLock()) {
return;
}
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
Sleep = false;
} else {
// Waiting for lock to become available. Add to waiters.
Desired = Expected | READ_WAITER_BIT;
Sleep = true;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Somehow read-locked and got a write lock!");
return;
}
FutexWaitForReadAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void unlock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Trying to write-unlock something not write-locked!");
// Remove the exclusive lock bit.
Desired = Expected & ~WRITE_OWNED_BIT;
// If no more writers, then make sure to clear the read-waiters bit as well.
if ((Desired & WRITE_WAITER_COUNT_MASK) == 0) {
Desired &= ~READ_WAITER_BIT;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
// If success, then `Expected` has old value. Containing `READ_WAITER_BIT` which was just masked off, and also `WRITE_WAITER_COUNT_MASK`.
if ((Expected & WRITE_WAITER_COUNT_MASK)) {
// Handle write-write handoff.
FutexWakeWriter();
} else if ((Expected & READ_WAITER_BIT)) {
// Handle write-reader handoff.
FutexWakeReaders();
}
}
void unlock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Trying to read-unlock something write-locked!");
LOGMAN_THROW_A_FMT((Expected & READ_OWNER_COUNT_MASK) != 0, "Trying to read-unlock something not read-locked!");
// Decrement the shared counter.
Desired = Expected - READ_OWNER_INCREMENT;
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Desired = AtomicFutex.fetch_sub(READ_OWNER_INCREMENT) - READ_OWNER_INCREMENT;
#endif
// Handle read->write handoff if there are any waiting writers, and no readers left.
if ((Desired & WRITE_WAITER_COUNT_MASK) && (Desired & READ_OWNER_COUNT_MASK) == 0) {
FutexWakeWriter();
}
}
bool try_lock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = 0;
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
// try to CAS immediately.
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
// Can race with other threads trying to lock shared!
bool try_lock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
// Exclusively owned or has a list of waiting owners. Can't pass.
if ((Expected & WRITE_OWNED_BIT) || (Expected & WRITE_WAITER_COUNT_MASK)) {
return false;
}
// Try to add reader.
uint32_t Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
// Uncontended mutex check
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
#if !defined(_WIN32)
// Initialize the internal mutex object to its default initializer state.
// Should only ever be used in the child process when a Linux fork() has occured.
void StealAndDropActiveLocks() {
Futex = 0;
}
#endif
private:
#if !defined(_WIN32)
void FutexWaitForWriteAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Read-lock waiting for writers to drain out.
void FutexWaitForReadAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
// Read-Lock or Write-lock unlocked, wake one writer.
// - Read->Write handoff.
// - Write->Write handoff.
void FutexWakeWriter() {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, 1, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Write-lock unlocked, wake read-locks waiting.
void FutexWakeReaders() {
// Wake all readers.
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, INT_MAX, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
#else
// Writers wait for the full 32-bit futex.
void FutexWaitForWriteAvailable(uint32_t Expected) {
WaitOnAddress(&Futex, &Expected, sizeof(Futex), INFINITE);
}
// Readers wait for Futex bits [31:16] to be zero.
void FutexWaitForReadAvailable(uint32_t Expected) {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
uint16_t smol_Expected = Expected >> 16;
WaitOnAddress(ReadWaiterAddress, &smol_Expected, sizeof(smol_Expected), INFINITE);
}
void FutexWakeWriter() {
WakeByAddressSingle(&Futex);
}
void FutexWakeReaders() {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
WakeByAddressAll(ReadWaiterAddress);
}
#endif
// Reuse the SpinWaitLock WFE implementations for read/write lock acquiring with WFE.
// Can't reuse the spin-lock directly as some bit-representations are different.
// WFE-write-lock is less likely to occur the more read-lock threads are participating. Can still occur so good to try.
// WFE-read-lock is actually quite likely to succeed.
// Return: true if the lock was acquired.
bool Attempt_WFE_WriteLock() {
#ifdef _M_ARM_64
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
while ((Now - Begin) < Duration) {
if (Expected == 0) {
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, ~0U, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
// Return: true if the lock was acquired.
bool Attempt_WFE_ReadLock() {
#ifdef _M_ARM_64
// Spin on a WFE for a short-amount of time, waiting for write-owned and writer-count to be zero.
// - Attempt to acquire read-lock at that point.
// - Don't add read-waiters bit on failure, return false.
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while ((Now - Begin) < Duration) {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, WRITE_OWNED_BIT | WRITE_WAITER_COUNT_MASK, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
constexpr static uint32_t WRITE_OWNED_BIT = 1U << 31;
constexpr static uint32_t READ_WAITER_BIT = 1U << 15;
constexpr static uint32_t WRITE_WAITER_OFFSET = 16;
constexpr static uint32_t WRITE_WAITER_INCREMENT = 1U << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_INCREMENT = 1;
// Count masks
constexpr static uint32_t WRITE_WAITER_COUNT_MASK = 0x7FFFU << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_COUNT_MASK = 0x7FFFU;
// Independent futex bit-set masks.
// Wait for readers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_READERS = 1U << 0;
// Wait for writers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_WRITERS = 1U << 1;
// Only spin on WFE for 0.01ms (10k ns).
constexpr static uint64_t CYCLECOUNT_DIVISOR = 1'000'000'000ULL / 10'000U;
// Layout:
// Bits[31]: Write-lock bit.
// Bits[30:16]: Write-waiter count.
// Bits[15]: Read-waiter bit.
// Bits[14:0]: Read-owner count.
uint32_t Futex {};
};
} // namespace FEXCore::Utils::WritePriorityMutex
+63 -33
View File
@@ -103,28 +103,25 @@ static inline std::optional<fextl::string> EnumParser(const ArrayPairType& EnumP
return fextl::fmt::format("{}", EnumMask);
}
namespace DefaultValues {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) extern const P(type) P(enum);
#define OPT_STR(group, enum, json, default) extern const std::string_view P(enum);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#include <FEXCore/Config/ConfigValues.inl>
using StringArrayType = fextl::list<fextl::string>;
namespace Type {
using StringArrayType = fextl::list<fextl::string>;
#define OPT_BASE(type, group, enum, json, default) using P(enum) = P(type);
#define OPT_STR(group, enum, json, default) using P(enum) = fextl::string;
#define OPT_STRARRAY(group, enum, json, default) using P(enum) = StringArrayType;
namespace detail {
template<ConfigOption Option>
struct ConfigOptionInfo;
#define DEFINE_METAINFO(type, enum, default) \
template<> \
struct ConfigOptionInfo<ConfigOption::CONFIG_##enum> { \
using Type = type; \
static auto Default() { \
extern default; \
return enum; \
} \
};
#define OPT_BASE(type, group, enum, json, default) DEFINE_METAINFO(type, enum, const type enum)
#define OPT_STR(group, enum, json, default) DEFINE_METAINFO(fextl::string, enum, const std::string_view enum)
#define OPT_STRARRAY(group, enum, json, default) DEFINE_METAINFO(StringArrayType, enum, const std::string_view enum)
#include <FEXCore/Config/ConfigValues.inl>
} // namespace Type
#define FEX_CONFIG_OPT(name, enum) \
FEXCore::Config::Value<FEXCore::Config::DefaultValues::Type::enum> name { \
FEXCore::Config::CONFIG_##enum, \
FEXCore::Config::DefaultValues::enum \
}
#undef P
} // namespace DefaultValues
} // namespace detail
FEX_DEFAULT_VISIBILITY void SetDataDirectory(std::string_view Path, bool Global);
FEX_DEFAULT_VISIBILITY void SetConfigDirectory(const std::string_view Path, bool Global);
@@ -135,8 +132,7 @@ FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigDirectory(bool Global);
FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigFileLocation(bool Global = false);
FEX_DEFAULT_VISIBILITY fextl::string GetApplicationConfig(const std::string_view Program, bool Global);
using LayerValue =
std::variant< fextl::string, DefaultValues::Type::StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerValue = std::variant< fextl::string, StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerOptions = fextl::unordered_map<ConfigOption, LayerValue>;
@@ -151,16 +147,16 @@ public:
return OptionMap.find(Option) != OptionMap.end();
}
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
std::optional<StringArrayType*> All(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
return std::nullopt;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
return &std::get<DefaultValues::Type::StringArrayType>(Value);
return &std::get<StringArrayType>(Value);
}
std::optional<fextl::string*> Get(ConfigOption Option) {
@@ -201,12 +197,12 @@ public:
auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
// If the option didn't exist as a StringArrayType yet, emplace it.
it = OptionMap.emplace(Option, DefaultValues::Type::StringArrayType {}).first;
it = OptionMap.emplace(Option, StringArrayType {}).first;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<DefaultValues::Type::StringArrayType>(Value).emplace_back(Data);
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<StringArrayType>(Value).emplace_back(Data);
}
void Erase(ConfigOption Option) {
@@ -236,7 +232,9 @@ FEX_DEFAULT_VISIBILITY fextl::string FindContainerPrefix();
FEX_DEFAULT_VISIBILITY void AddLayer(fextl::unique_ptr<FEXCore::Config::Layer> _Layer);
FEX_DEFAULT_VISIBILITY bool Exists(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<StringArrayType*> All(ConfigOption Option);
template<typename T>
FEX_DEFAULT_VISIBILITY std::optional<T> GetConv(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<fextl::string*> Get(ConfigOption Option);
FEX_DEFAULT_VISIBILITY void Set(ConfigOption Option, std::string_view Data);
FEX_DEFAULT_VISIBILITY void Erase(ConfigOption Option);
@@ -271,18 +269,18 @@ public:
return ValueData;
}
Value(T Value) requires (!std::is_same_v<T, DefaultValues::Type::StringArrayType>)
Value(T Value) requires (!std::is_same_v<T, StringArrayType>)
{
ValueData = std::move(Value);
}
// Array value types.
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, StringArrayType>)
{
GetListIfExists(Option, &ValueData);
}
DefaultValues::Type::StringArrayType& All() requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
StringArrayType& All() requires (std::is_same_v<T, StringArrayType>)
{
return ValueData;
}
@@ -293,6 +291,38 @@ private:
static T GetIfExists(FEXCore::Config::ConfigOption Option, T Default);
static T GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
};
/**
* Wrapper around Value that automatically picks the default for the given ConfigOption
*/
template<ConfigOption Option>
struct FEX_DEFAULT_VISIBILITY Getter : public Value<typename detail::ConfigOptionInfo<Option>::Type> {
using OptionInfo = detail::ConfigOptionInfo<Option>;
Getter()
: Value<typename OptionInfo::Type> {Option, OptionInfo::Default()} {}
};
/**
* Helper for reading a config value with caching.
*
* Typically this is used to declare class members so that the value is read
* on construction of the parent.
*/
#define FEX_CONFIG_OPT(name, enum) FEXCore::Config::Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> name {}
#define OPT_BASE(type, group, enum, json, default) \
/** \
* Helper for reading a config value. \
* \
* In contrast to FEX_CONFIG_OPT, this can be used in arbitrary expressions, \
* at the expense of not caching the value. Use Getter instead if the value \
* is read frequently. \
*/ \
inline auto Get_##enum() { \
return Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> {}; \
}
#include <FEXCore/Config/ConfigValues.inl>
} // namespace FEXCore::Config
+136 -1
View File
@@ -1,10 +1,20 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/functional.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <atomic>
#include <cstdint>
#include <mutex>
#include <optional>
#include <shared_mutex>
#include <span>
#include <unistd.h>
namespace FEXCore {
@@ -20,8 +30,14 @@ namespace HLE {
struct ExecutableFileInfo {
~ExecutableFileInfo();
#if __clang_major__ < 16
// Workaround for broken aggregate-initialization with std::piecewise_construct
ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap>, uint64_t, fextl::string);
ExecutableFileInfo() = default;
#endif
fextl::unique_ptr<HLE::SourcecodeMap> SourcecodeMap;
fextl::string FileId;
uint64_t FileId = 0;
fextl::string Filename;
};
@@ -33,10 +49,129 @@ struct ExecutableFileSectionInfo {
uintptr_t FileStartVA;
};
using CodeMapFileId = uint64_t;
/**
* Code maps capture information required for offline code cache generation
* and are written to disk during execution of FEX.
*
* Almost all CodeMap data will be an Entry that indicates blocks to be
* compiled for cache generation. The reserved value `LoadExternalLibrary`
* indicates that an instance of ExternalLibraryInfo follows (the entry data
* itself should be skipped in that case).
*/
struct CodeMap {
// Describes the location of an entry block compiled during execution
struct FEX_PACKED Entry {
CodeMapFileId FileId;
uint32_t BlockOffset;
};
// Describes an external library referenced during execution
struct ExternalLibraryInfo {
CodeMapFileId ExternalFileId;
// null-terminated file path; EITHER relative to the main executable OR an absolute path OR starting with a magic identifier:
// - WINE/: Path to Wine/Proton installation
// - WINEPREFIX/: Path to Wine/Proton prefix
// - SLR/: Path to Steam Linux Runtime
// At runtime, FEX will always dump absolute paths
char Path[];
// Followed by padding to a 4 byte boundary
};
// Followed by ExternalLibraryInfo
static constexpr Entry LoadExternalLibrary = {0xffff'ffff'ffff'ffff, 0xffff'ffff};
struct FEX_PACKED SetExecutableFileId {
Entry Marker = {0xffff'ffff'ffff'ffff, 0xffff'fffe};
CodeMapFileId ExecutableFileId;
};
struct ParsedContents {
fextl::string Filename;
fextl::set<uint64_t> Blocks;
bool IsExecutable = false;
};
// Follows scheme fileid[-nomb]
// The nomb ("no multiblock") suffix signifies that the code map is for use without multiblock, only.
static fextl::string GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix);
static fextl::map<CodeMapFileId, ParsedContents> ParseCodeMap(std::ifstream& File);
};
struct CodeMapOpener {
virtual ~CodeMapOpener() = default;
virtual int OpenCodeMapFile() = 0;
};
class CodeMapWriter {
public:
CodeMapWriter(CodeMapOpener&, bool OpenEagerly = false);
~CodeMapWriter();
// Checks if writing is enabled. Calls to this functions may also be interpreted as signals that writes are about to happen
bool IsWriteEnabled(const ExecutableFileSectionInfo&);
void ResetAfterFork() {
if (CodeMapFD.value_or(-1) != -1) {
close(CodeMapFD.value());
CodeMapFD.reset();
}
BufferOffset = 0;
KnownFileIds.clear();
}
bool IsBackingFD(int FD) const {
if (FD == CodeMapFD) {
LogMan::Msg::DFmt("Hiding directory entry for code map FD");
return true;
}
return false;
}
void AppendBlock(const FEXCore::ExecutableFileSectionInfo&, uint64_t Entry);
void AppendLibraryLoad(const FEXCore::ExecutableFileInfo&);
void AppendSetMainExecutable(const FEXCore::ExecutableFileInfo&);
// Thread-safely commit any pending data to disk
void Flush(size_t Offset);
private:
// Queues data into an internal ring buffer.
// Call Flush() to commit the data to disk.
void AppendData(std::span<const std::byte> Data);
// Commit given data range to disk
void Flush(size_t Offset, std::unique_lock<std::shared_mutex>&);
std::shared_mutex Mutex;
fextl::vector<std::byte> Buffer;
std::atomic<size_t> BufferOffset {0};
fextl::set<CodeMapFileId> KnownFileIds;
// std::nullopt: We haven't requested a CodeMapFD yet
// value is -1: We requested a CodeMapFD but FEXServer told us not to write any data
// other values: Code map writing is active
std::optional<int> CodeMapFD;
CodeMapOpener& FileOpener;
};
class AbstractCodeCache {
public:
virtual ~AbstractCodeCache() = default;
/**
* Computes a unique identifier for the referenced binary file to be used for
* generating the code map.
* This identifier is independent of FEX build/runtime configuration and
* stable across FEX updates.
*/
virtual uint64_t ComputeCodeMapId(std::string_view Filename, int FD) = 0;
/**
* Loads a code cache from mapped memory and appends it to the current Core state.
* TODO: Optionally recompiles all contained code blocks at runtime for validation.
+5 -5
View File
@@ -42,9 +42,6 @@ enum OperatingMode {
using CodeRangeInvalidationFn = std::function<void(uint64_t start, uint64_t Length)>;
// Nested vector of guest block entrypoints
using InvalidatedEntryAccumulator = fextl::vector<fextl::vector<uint64_t>>;
using CustomIREntrypointHandler = std::function<void(uintptr_t Entrypoint, IR::IREmitter*)>;
using ExitHandler = std::function<void(Core::InternalThreadState* Thread)>;
@@ -139,10 +136,13 @@ public:
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) = 0;
virtual AbstractCodeCache& GetCodeCache() = 0;
virtual void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter>) = 0;
virtual void FlushAndCloseCodeMap() = 0;
FEX_DEFAULT_VISIBILITY virtual void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(
FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void
InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual void
-5
View File
@@ -353,7 +353,6 @@ struct JITPointers {
uint64_t ExitFunctionLinker {};
uint64_t ThreadStopHandlerSpillSRA {};
uint64_t ThreadPauseHandlerSpillSRA {};
uint64_t UnimplementedInstructionHandler {};
uint64_t GuestSignal_SIGILL {};
uint64_t GuestSignal_SIGTRAP {};
uint64_t GuestSignal_SIGSEGV {};
@@ -371,8 +370,6 @@ struct JITPointers {
// Process specific
uint64_t LUDIV {};
uint64_t LDIV {};
uint64_t LUREM {};
uint64_t LREM {};
// Thread Specific
@@ -381,8 +378,6 @@ struct JITPointers {
* @{ */
uint64_t LUDIVHandler {};
uint64_t LDIVHandler {};
uint64_t LUREMHandler {};
uint64_t LREMHandler {};
/** @} */
} AArch64;
@@ -67,6 +67,10 @@ public:
return Config;
}
virtual uintptr_t GetThunkCallbackRET() const {
return 0;
}
protected:
SignalDelegatorConfig Config;
};
@@ -4,6 +4,7 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/vector.h>
@@ -118,6 +119,10 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
// The low address of the call-ret stack allocation (not including guard pages)
void* CallRetStackBase {};
uintptr_t JITGuardPage {};
uint64_t JITGuardOverflowArgument {};
FEXCore::UncheckedLongJump::JumpBuf RestartJump;
// BaseFrameState should always be at the end, directly before the interrupt fault page
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState {};
+2 -9
View File
@@ -3,6 +3,7 @@
#include <FEXCore/Utils/EnumOperators.h>
#include <compare>
#include <cstdint>
#include <cstring>
@@ -76,15 +77,7 @@ enum IndexNamedVectorConstant : uint8_t {
struct SHA256Sum final {
uint8_t data[32];
[[nodiscard]]
bool operator<(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) < 0;
}
[[nodiscard]]
bool operator==(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) == 0;
}
[[nodiscard]] auto operator<=>(const SHA256Sum&) const noexcept = default;
};
typedef void ThunkedFunction(void* ArgsRv);
+39
View File
@@ -0,0 +1,39 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <cstdint>
// Reimplementation of longjmp without glibc fortification checks.
// This is useful when false positives need to be avoided or when using
// a libc implementation that does not implement std::longjmp.
namespace FEXCore::UncheckedLongJump {
// JumpBuf definition needs to be public because the frontend needs to understand it.
#if defined(_M_ARM_64)
struct JumpBuf {
// All the registers that are required by AAPCS64 to save.
// GPRs
// X19, X20, X21, X22,
// X23, X24, X25, X26,
// X27, X28, X29, X30,
//
// Lower 64-bits:
// V8, V9, V10, V11,
// V12, V13, V14, V15,
//
// SP,
uint64_t Registers[21];
};
#else
struct JumpBuf {
// Registers to preserve
// RBX, RSP, RBP, R12, R13, R14, R15,
// <return address>
uint64_t Registers[8];
};
#endif
[[nodiscard]] FEX_DEFAULT_VISIBILITY uint64_t SetJump(JumpBuf& Buffer);
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(const JumpBuf& Buffer, uint64_t Value);
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC);
} // namespace FEXCore::UncheckedLongJump
+1
View File
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <cstddef>
#include <cstdint>
#ifdef _M_X86_64
@@ -28,4 +28,19 @@ inline fextl::string Trim(fextl::string String, std::string_view TrimTokens = "
return RightTrim(LeftTrim(std::move(String), TrimTokens), TrimTokens);
}
inline fextl::string& ReplaceAllInPlace(fextl::string& Str, std::string_view Token, std::string_view New) {
const auto OriginalTokenSize = Token.size();
const auto NewTokenSize = New.size();
size_t TokenPos {};
auto TokenIter = Str.find(Token, TokenPos);
while (TokenIter != Str.npos) {
Str.replace(TokenIter, OriginalTokenSize, New);
TokenPos += NewTokenSize;
TokenIter = Str.find(Token, TokenPos);
}
return Str;
}
} // namespace FEXCore::StringUtils
@@ -3,6 +3,8 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/list.h>
#include <atomic>
@@ -37,12 +39,6 @@ namespace FEXCore::Utils {
*/
class IntrusivePooledAllocator {
public:
template<typename T>
struct AllocationInfo {
T Ptr;
size_t Size;
};
struct MemoryBuffer;
/**
* @brief Container for tracking the buffers
@@ -403,6 +399,42 @@ private:
const char* Name {};
};
/**
* @brief Thread pool allocator that allocates and frees objects that uses mmap, with a guard page.
*
* The last page of the size provided has the guard.
*/
class PooledAllocatorVirtualWithGuard final : public IntrusivePooledAllocator {
public:
PooledAllocatorVirtualWithGuard() = default;
PooledAllocatorVirtualWithGuard(const char* Name)
: Name {Name} {}
virtual ~PooledAllocatorVirtualWithGuard() {
FreeAllBuffers();
}
private:
void* Alloc(size_t Size) override {
auto Ptr = FEXCore::Allocator::VirtualAlloc(Size);
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
if (Name) {
FEXCore::Allocator::VirtualName(Name, Ptr, Size);
}
return Ptr;
}
void Free(void* Ptr, size_t Size) override {
FEXCore::Allocator::VirtualFree(Ptr, Size);
}
const char* Name {};
};
/**
* @brief Wrapper around the pool allocator for delayed pool reclaiming
*
@@ -460,6 +492,11 @@ public:
UnclaimBuffer();
}
struct AllocationInfo {
Type Ptr;
size_t Size;
};
/**
* @brief Return the owned buffer or allocate another one from the `Allocator`
*
@@ -468,9 +505,9 @@ public:
*
* @param NewSize Optional new size for managed data
*
* @return object of type `Type` allocated within the selected buffer
* @return A usable pointer of type `Type` and the size of the backing store.
*/
Type ReownOrClaimBuffer(std::optional<size_t> NewSize = std::nullopt) {
AllocationInfo ReownOrClaimBufferWithSize(std::optional<size_t> NewSize = std::nullopt) {
// Check if we can cheaply re-own a previous buffer
std::optional Buffer =
IntrusivePooledAllocator::IsClientBufferOwned(ClientOwnedFlag) ? Info : ThreadAllocator.TryToReownBuffer(Info, Size, &ClientOwnedFlag);
@@ -493,7 +530,14 @@ public:
// Leaving this here for future excavation that will definitely occur here
// memset((*Info)->Ptr, 0, Size);
return reinterpret_cast<Type>((*Info)->Ptr);
return {
.Ptr = reinterpret_cast<Type>((*Info)->Ptr),
.Size = (*Info)->Size,
};
}
Type ReownOrClaimBuffer(std::optional<size_t> NewSize = std::nullopt) {
return ReownOrClaimBufferWithSize(NewSize).Ptr;
}
/**
+10
View File
@@ -0,0 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/allocator.h>
#include <tsl/robin_set.h>
namespace fextl {
template<class Key, class Hash = std::hash<Key>, class KeyEqual = std::equal_to<Key>, class Allocator = fextl::FEXAlloc<Key>>
using robin_set = tsl::robin_set<Key, Hash, KeyEqual, Allocator>;
}
+66
View File
@@ -0,0 +1,66 @@
// SPDX-License-Identifier: MIT
#include <catch2/catch_test_macros.hpp>
#include <catch2/generators/catch_generators_range.hpp>
#include "Utils/Allocator/HostAllocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <sys/mman.h>
template<typename T>
bool HasSyscallError(T Result) {
constexpr uint64_t MAX_ERRNO = 0xFFFF'FFFF'FFFF'0001ULL;
return reinterpret_cast<uint64_t>(Result) >= MAX_ERRNO;
}
TEST_CASE("Allocator - Fixed replacement") {
const auto RegionSize = 128 * 1024 * 1024;
fextl::vector<FEXCore::Allocator::MemoryRegion> MemoryRegions {};
for (size_t i = 0; i < 2; ++i) {
auto Ptr = mmap(nullptr, RegionSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
MemoryRegions.emplace_back(FEXCore::Allocator::MemoryRegion {
.Ptr = Ptr,
.Size = RegionSize,
});
}
auto Allocator = Alloc::OSAllocator::Create64BitAllocatorWithRegions(MemoryRegions);
auto Base = Allocator->Mmap(nullptr, 4096, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(!HasSyscallError(Base));
// Allocate perfectly overlapping pages. Allocate as many pages as the region.
// FEX had a bug where the allocator could run out of memory with MAP_FIXED.
for (size_t i = 0; i < (RegionSize / 4096); ++i) {
auto NewBase = Allocator->Mmap(Base, 4096, PROT_NONE, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(Base == NewBase);
}
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Allocator));
}
TEST_CASE("Allocator - Non-Fit") {
const auto RegionSize = 128 * 1024 * 1024;
fextl::vector<FEXCore::Allocator::MemoryRegion> MemoryRegions {};
for (size_t i = 0; i < 2; ++i) {
auto Ptr = mmap(nullptr, RegionSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
MemoryRegions.emplace_back(FEXCore::Allocator::MemoryRegion {
.Ptr = Ptr,
.Size = RegionSize,
});
}
auto Allocator = Alloc::OSAllocator::Create64BitAllocatorWithRegions(MemoryRegions);
auto Base = Allocator->Mmap(nullptr, RegionSize / 4, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(!HasSyscallError(Base));
// Try to allocate within the whole VMA size minus a small amount.
// FEX had a bug where if the allocation fit within a VMA region, it would try and allocate past the end without checking.
// Only occurred when `MAP_FIXED` was used.
auto NewBase = Allocator->Mmap(Base, RegionSize - (4096 * 64), PROT_NONE, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
// Must either fit in the VMA region, or fail.
// - If it matches previous allocation, then it fit in the VMA region.
// - This can happen if FEX's allocator gains support for VMA merging.
// - If it errors, then it doesn't fit in the VMA region.
REQUIRE((NewBase == Base || HasSyscallError(NewBase)));
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Allocator));
}
+23
View File
@@ -3,6 +3,7 @@
#include <catch2/generators/catch_generators_range.hpp>
#include "Utils/Allocator/FlexBitSet.h"
#include <sys/mman.h>
TEST_CASE("FlexBitSet - Sizing") {
// Ensure that FlexBitSet sizing is correct.
@@ -40,3 +41,25 @@ TEST_CASE("FlexBitSet - Sizing") {
CHECK(FEXCore::FlexBitSet<uint32_t>::SizeInBits(sizeof(uint32_t) * 8) == sizeof(uint32_t) * 8);
CHECK(FEXCore::FlexBitSet<uint64_t>::SizeInBits(sizeof(uint64_t) * 8) == sizeof(uint64_t) * 8);
}
TEST_CASE("FlexBitSet - Limit") {
// Ensure that the FlexBitSet doesn't read past the limits, and returns correct indexes.
const auto Size = 4096 * 3;
auto Ptr = mmap(nullptr, Size, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
auto PtrMiddle = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(Ptr) + 4096);
REQUIRE(mprotect(PtrMiddle, 4096, PROT_READ | PROT_WRITE) != -1);
using ElementType = uint8_t;
const size_t NumElements = 4096 * 8;
auto FlexBit = reinterpret_cast<FEXCore::FlexBitSet<ElementType>*>(PtrMiddle);
for (size_t i = 0; i < NumElements; ++i) {
auto Result = FlexBit->ForwardScanForRange<true>(i, 1, NumElements);
CHECK(Result.FoundElement == i);
}
for (size_t i = 0; i < NumElements; ++i) {
auto Result = FlexBit->BackwardScanForRange<true>(i, 1, 0);
CHECK(Result.FoundElement == i);
}
}
+58 -52
View File
@@ -9,17 +9,17 @@ using namespace ARMEmitter;
TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
adr(Reg::r30, &Label);
(void)adr(Reg::r30, &Label);
CHECK(DisassembleEncoding(1) == 0x10fffffe);
}
{
ForwardLabel Label;
adr(Reg::r30, &Label);
Bind(&Label);
(void)adr(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x1000003e);
@@ -27,17 +27,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
adr(Reg::r30, &Label);
(void)adr(Reg::r30, &Label);
CHECK(DisassembleEncoding(1) == 0x10fffffe);
}
{
BiDirectionalLabel Label;
adr(Reg::r30, &Label);
Bind(&Label);
(void)adr(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x1000003e);
@@ -45,42 +45,42 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
adrp(Reg::r30, &Label);
(void)adrp(Reg::r30, &Label);
CHECK(DisassembleEncoding(1) == 0x9000001e);
}
{
ForwardLabel Label;
adrp(Reg::r30, &Label);
(void)adrp(Reg::r30, &Label);
// Move label a page away
for (size_t i = 0; i < 1023; ++i) {
nop();
}
Bind(&Label);
(void)Bind(&Label);
CHECK(DisassembleEncoding(0) == 0xb000001e);
}
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
adrp(Reg::r30, &Label);
(void)adrp(Reg::r30, &Label);
CHECK(DisassembleEncoding(1) == 0x9000001e);
}
{
BiDirectionalLabel Label;
adrp(Reg::r30, &Label);
(void)adrp(Reg::r30, &Label);
// Move label a page away
for (size_t i = 0; i < 1023; ++i) {
nop();
}
Bind(&Label);
(void)Bind(&Label);
CHECK(DisassembleEncoding(0) == 0xb000001e);
}
@@ -88,47 +88,49 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
{
// Will generate adr.
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
CHECK(DisassembleEncoding(1) == 0x10fffffe);
}
{
// Will generate nop + adr.
// Will generate nop + nop + adr.
ForwardLabel Label;
LongAddressGen(Reg::r30, &Label);
Bind(&Label);
(void)LongAddressGen(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x1000003e);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x1000003e);
}
{
// Will generate adr.
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
CHECK(DisassembleEncoding(1) == 0x10fffffe);
}
{
// Will generate nop + adr.
// Will generate nop + nop + adr.
BiDirectionalLabel Label;
LongAddressGen(Reg::r30, &Label);
Bind(&Label);
(void)LongAddressGen(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x1000003e);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x1000003e);
}
{
// Will generate adrp.
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
// Move adrp 1MB away.
@@ -136,51 +138,53 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
nop();
}
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
nop();
CHECK(DisassembleEncoding(262145) == 0x90fff81e);
CHECK(DisassembleEncoding(262146) == 0xd503201f);
}
{
// Will generate nop + adrp.
// Will generate nop + nop + adrp.
ForwardLabel Label;
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, and then aligned to a page.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 2); ++i) {
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 3); ++i) {
nop();
}
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x9000081e);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x9000081e);
}
{
// Will generate adrp + add.
// Will generate nop + adrp + add.
ForwardLabel Label;
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, plus one instruction.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 1); ++i) {
nop();
}
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb000081e);
CHECK(DisassembleEncoding(1) == 0x910013de);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xb000081e);
CHECK(DisassembleEncoding(2) == 0x910013de);
}
{
// Will generate adrp.
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
// Move adrp 1MB away.
@@ -188,44 +192,46 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
nop();
}
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
nop();
CHECK(DisassembleEncoding(262145) == 0x90fff81e);
CHECK(DisassembleEncoding(262146) == 0xd503201f);
}
{
// Will generate nop + adrp.
// Will generate nop + nop + adrp.
BiDirectionalLabel Label;
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, and then aligned to a page.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 2); ++i) {
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 3); ++i) {
nop();
}
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x9000081e);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x9000081e);
}
{
// Will generate adrp + add.
// Will generate nop + adrp + add.
BiDirectionalLabel Label;
LongAddressGen(Reg::r30, &Label);
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, plus one instruction.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 1); ++i) {
nop();
}
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb000081e);
CHECK(DisassembleEncoding(1) == 0x910013de);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xb000081e);
CHECK(DisassembleEncoding(2) == 0x910013de);
}
}
TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: Add/subtract immediate") {
+96 -96
View File
@@ -9,17 +9,17 @@ using namespace ARMEmitter;
TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Conditional branch immediate") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
b(Condition::CC_PL, &Label);
(void)b(Condition::CC_PL, &Label);
CHECK(DisassembleEncoding(1) == 0x54ffffe5);
}
{
ForwardLabel Label;
b(Condition::CC_PL, &Label);
Bind(&Label);
(void)b(Condition::CC_PL, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x54000025);
@@ -27,17 +27,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Conditional branch immediat
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
b(Condition::CC_PL, &Label);
(void)b(Condition::CC_PL, &Label);
CHECK(DisassembleEncoding(1) == 0x54ffffe5);
}
{
BiDirectionalLabel Label;
b(Condition::CC_PL, &Label);
Bind(&Label);
(void)b(Condition::CC_PL, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x54000025);
@@ -46,17 +46,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Conditional branch immediat
TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Branch consistent conditional") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
bc(Condition::CC_PL, &Label);
(void)bc(Condition::CC_PL, &Label);
CHECK(DisassembleEncoding(1) == 0x54fffff5);
}
{
ForwardLabel Label;
bc(Condition::CC_PL, &Label);
Bind(&Label);
(void)bc(Condition::CC_PL, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x54000035);
@@ -64,17 +64,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Branch consistent condition
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
bc(Condition::CC_PL, &Label);
(void)bc(Condition::CC_PL, &Label);
CHECK(DisassembleEncoding(1) == 0x54fffff5);
}
{
BiDirectionalLabel Label;
bc(Condition::CC_PL, &Label);
Bind(&Label);
(void)bc(Condition::CC_PL, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x54000035);
@@ -89,17 +89,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Unconditional branch regist
TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Unconditional branch immediate") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
b(&Label);
(void)b(&Label);
CHECK(DisassembleEncoding(1) == 0x17ffffff);
}
{
ForwardLabel Label;
b(&Label);
Bind(&Label);
(void)b(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x14000001);
@@ -107,17 +107,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Unconditional branch immedi
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
b(&Label);
(void)b(&Label);
CHECK(DisassembleEncoding(1) == 0x17ffffff);
}
{
BiDirectionalLabel Label;
b(&Label);
Bind(&Label);
(void)b(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x14000001);
@@ -125,17 +125,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Unconditional branch immedi
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
bl(&Label);
(void)bl(&Label);
CHECK(DisassembleEncoding(1) == 0x97ffffff);
}
{
ForwardLabel Label;
bl(&Label);
Bind(&Label);
(void)bl(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x94000001);
@@ -143,17 +143,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Unconditional branch immedi
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
bl(&Label);
(void)bl(&Label);
CHECK(DisassembleEncoding(1) == 0x97ffffff);
}
{
BiDirectionalLabel Label;
bl(&Label);
Bind(&Label);
(void)bl(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x94000001);
@@ -162,17 +162,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Unconditional branch immedi
TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbz(Size::i32Bit, Reg::r29, &Label);
(void)cbz(Size::i32Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0x34fffffd);
}
{
ForwardLabel Label;
cbz(Size::i32Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbz(Size::i32Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3400003d);
@@ -180,17 +180,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbz(Size::i32Bit, Reg::r29, &Label);
(void)cbz(Size::i32Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0x34fffffd);
}
{
BiDirectionalLabel Label;
cbz(Size::i32Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbz(Size::i32Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3400003d);
@@ -198,17 +198,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbz(Size::i64Bit, Reg::r29, &Label);
(void)cbz(Size::i64Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0xb4fffffd);
}
{
ForwardLabel Label;
cbz(Size::i64Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbz(Size::i64Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb400003d);
@@ -216,17 +216,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbz(Size::i64Bit, Reg::r29, &Label);
(void)cbz(Size::i64Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0xb4fffffd);
}
{
BiDirectionalLabel Label;
cbz(Size::i64Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbz(Size::i64Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb400003d);
@@ -234,17 +234,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbnz(Size::i32Bit, Reg::r29, &Label);
(void)cbnz(Size::i32Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0x35fffffd);
}
{
ForwardLabel Label;
cbnz(Size::i32Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbnz(Size::i32Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3500003d);
@@ -252,17 +252,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbnz(Size::i32Bit, Reg::r29, &Label);
(void)cbnz(Size::i32Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0x35fffffd);
}
{
BiDirectionalLabel Label;
cbnz(Size::i32Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbnz(Size::i32Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3500003d);
@@ -270,17 +270,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbnz(Size::i64Bit, Reg::r29, &Label);
(void)cbnz(Size::i64Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0xb5fffffd);
}
{
ForwardLabel Label;
cbnz(Size::i64Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbnz(Size::i64Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb500003d);
@@ -288,17 +288,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
cbnz(Size::i64Bit, Reg::r29, &Label);
(void)cbnz(Size::i64Bit, Reg::r29, &Label);
CHECK(DisassembleEncoding(1) == 0xb5fffffd);
}
{
BiDirectionalLabel Label;
cbnz(Size::i64Bit, Reg::r29, &Label);
Bind(&Label);
(void)cbnz(Size::i64Bit, Reg::r29, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb500003d);
@@ -307,17 +307,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Compare and branch") {
TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbz(Reg::r29, 0, &Label);
(void)tbz(Reg::r29, 0, &Label);
CHECK(DisassembleEncoding(1) == 0x3607fffd);
}
{
ForwardLabel Label;
tbz(Reg::r29, 0, &Label);
Bind(&Label);
(void)tbz(Reg::r29, 0, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3600003d);
@@ -325,17 +325,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate")
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbz(Reg::r29, 0, &Label);
(void)tbz(Reg::r29, 0, &Label);
CHECK(DisassembleEncoding(1) == 0x3607fffd);
}
{
BiDirectionalLabel Label;
tbz(Reg::r29, 0, &Label);
Bind(&Label);
(void)tbz(Reg::r29, 0, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3600003d);
@@ -343,17 +343,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbz(Reg::r29, 63, &Label);
(void)tbz(Reg::r29, 63, &Label);
CHECK(DisassembleEncoding(1) == 0xb6fffffd);
}
{
ForwardLabel Label;
tbz(Reg::r29, 63, &Label);
Bind(&Label);
(void)tbz(Reg::r29, 63, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb6f8003d);
@@ -361,17 +361,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate")
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbz(Reg::r29, 63, &Label);
(void)tbz(Reg::r29, 63, &Label);
CHECK(DisassembleEncoding(1) == 0xb6fffffd);
}
{
BiDirectionalLabel Label;
tbz(Reg::r29, 63, &Label);
Bind(&Label);
(void)tbz(Reg::r29, 63, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb6f8003d);
@@ -379,17 +379,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbnz(Reg::r29, 0, &Label);
(void)tbnz(Reg::r29, 0, &Label);
CHECK(DisassembleEncoding(1) == 0x3707fffd);
}
{
ForwardLabel Label;
tbnz(Reg::r29, 0, &Label);
Bind(&Label);
(void)tbnz(Reg::r29, 0, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3700003d);
@@ -397,17 +397,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate")
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbnz(Reg::r29, 0, &Label);
(void)tbnz(Reg::r29, 0, &Label);
CHECK(DisassembleEncoding(1) == 0x3707fffd);
}
{
BiDirectionalLabel Label;
tbnz(Reg::r29, 0, &Label);
Bind(&Label);
(void)tbnz(Reg::r29, 0, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x3700003d);
@@ -415,17 +415,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbnz(Reg::r29, 63, &Label);
(void)tbnz(Reg::r29, 63, &Label);
CHECK(DisassembleEncoding(1) == 0xb7fffffd);
}
{
ForwardLabel Label;
tbnz(Reg::r29, 63, &Label);
Bind(&Label);
(void)tbnz(Reg::r29, 63, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb7f8003d);
@@ -433,17 +433,17 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Branch: Test and branch immediate")
{
BiDirectionalLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
tbnz(Reg::r29, 63, &Label);
(void)tbnz(Reg::r29, 63, &Label);
CHECK(DisassembleEncoding(1) == 0xb7fffffd);
}
{
BiDirectionalLabel Label;
tbnz(Reg::r29, 63, &Label);
Bind(&Label);
(void)tbnz(Reg::r29, 63, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb7f8003d);
+14 -14
View File
@@ -1323,7 +1323,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: LDAPR/STLR unscaled imme
TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal") {
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
ldr(WReg::w30, &Label);
@@ -1332,7 +1332,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
ldr(SReg::s30, &Label);
@@ -1341,7 +1341,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
ldr(XReg::x30, &Label);
@@ -1350,7 +1350,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
ldr(DReg::d30, &Label);
@@ -1359,7 +1359,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
ldrsw(XReg::x30, &Label);
@@ -1368,7 +1368,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
ldr(QReg::q30, &Label);
@@ -1377,7 +1377,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
BackwardLabel Label;
Bind(&Label);
(void)Bind(&Label);
dc32(0);
prfm(Prefetch::PLDL1KEEP, &Label);
@@ -1387,7 +1387,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
ForwardLabel Label;
ldr(WReg::w30, &Label);
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x1800003e);
@@ -1396,7 +1396,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
ForwardLabel Label;
ldr(SReg::s30, &Label);
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x1c00003e);
@@ -1405,7 +1405,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
ForwardLabel Label;
ldr(XReg::x30, &Label);
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x5800003e);
@@ -1414,7 +1414,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
ForwardLabel Label;
ldr(DReg::d30, &Label);
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x5c00003e);
@@ -1423,7 +1423,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
ForwardLabel Label;
ldrsw(XReg::x30, &Label);
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x9800003e);
@@ -1432,7 +1432,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
ForwardLabel Label;
ldr(QReg::q30, &Label);
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0x9c00003e);
@@ -1441,7 +1441,7 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: Loadstore: Load register literal")
{
ForwardLabel Label;
prfm(Prefetch::PLDL1KEEP, &Label);
Bind(&Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd8000020);
+1 -2
View File
@@ -1,6 +1,5 @@
#!/usr/bin/python3
import xxhash
import hashlib
import sys
import os
import shutil
@@ -188,5 +187,5 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
-3
View File
@@ -1,8 +1,5 @@
#!/usr/bin/python3
import os
import subprocess
import sys
import tempfile
import platform
def ListContainsRequired(Features, RequiredFeatures):
+2 -2
View File
@@ -4,7 +4,7 @@ from clang.cindex import CursorKind
from clang.cindex import TypeKind
from clang.cindex import TranslationUnit
import sys
from dataclasses import dataclass, field
from dataclasses import dataclass
import subprocess
import logging
logger = logging.getLogger()
@@ -548,5 +548,5 @@ def main():
PrintFunctionDecls()
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+3 -3
View File
@@ -9,19 +9,19 @@ for fileid in ~/.fex-emu/aotir/*.path; do
else
args="$args --no-abilocalflags"
fi
if [ "${fileid: -7 : 1}" == "T" ]; then
args="$args --tsoenabled"
else
args="$args --no-tsoenabled"
fi
if [ "${fileid: -8 : 1}" == "S" ]; then
args="$args --smc=full"
else
args="$args --smc=mman"
fi
if [ -f "${fileid%.path}.aotir" ]; then
echo "`basename $fileid` has already been generated"
else
+2 -2
View File
@@ -1,5 +1,5 @@
#!/usr/bin/python3
from dataclasses import dataclass, field
from dataclasses import dataclass
import math
import sys
import logging
@@ -282,5 +282,5 @@ def main():
ExportCommonSyscallDefines()
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+26 -21
View File
@@ -2,7 +2,6 @@
import os
import subprocess
import sys
import tempfile
import re
_Arch = None
@@ -203,10 +202,26 @@ def UpdatePPA():
return DidUpdate
def CheckAndInstallPackageUpdates():
PackagesToInstall = GetPackagesToInstall()
def InstallPackages(PackagesToInstall):
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + PackagesToInstall)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages updated")
else:
print("Packages failed to update")
return DidInstall
def CheckAndInstallPackageUpdates(PackagesToInstall, InstallIfNotFound=False):
for Package in PackagesToInstall[:]:
UpgradableStatus = subprocess.check_output(["apt", "list", "--upgradable", Package]).decode("utf-8")
UpgradableStatus = subprocess.check_output(["apt", "list", "--upgradable", Package], stderr=None).decode("utf-8")
Found = False
for Line in UpgradableStatus.split("\n"):
# If the package exists to be upgraded then it will appear in this list
@@ -221,28 +236,14 @@ def CheckAndInstallPackageUpdates():
if Package in Line and "upgradable" in Line:
Found = True
if Found == False:
if InstallIfNotFound == False and Found == False:
PackagesToInstall.remove(Package)
if len(PackagesToInstall) > 0:
print ("Found updates for packages: {}".format(PackagesToInstall))
print ("This bit may ask for your password")
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + PackagesToInstall)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages updated")
else:
print("Packages failed to update")
return DidInstall
return InstallPackages(PackagesToInstall)
return True
@@ -354,10 +355,14 @@ def main():
if not UpdatePPA():
print ("apt sources failed to update. Not continuing")
ExitWithStatus(-1)
if not CheckAndInstallPackageUpdates():
if not CheckAndInstallPackageUpdates(GetPackagesToInstall()):
print ("apt packages failed to update. Not continuing")
ExitWithStatus(-1)
else:
if not CheckAndInstallPackageUpdates(["software-properties-common"], True):
print ("software-properties-common package failed to update. Not continuing")
ExitWithStatus(-1)
if not InstallPPA():
print ("PPA failed to install. Not continuing")
ExitWithStatus(-1)
+2 -2
View File
@@ -1,6 +1,6 @@
#!/usr/bin/python3
import base64
from dataclasses import dataclass, field
from dataclasses import dataclass
from enum import Flag
import json
import struct
@@ -254,5 +254,5 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+2 -2
View File
@@ -4,7 +4,7 @@ from clang.cindex import CursorKind
from clang.cindex import TypeKind
from clang.cindex import TranslationUnit
import sys
from dataclasses import dataclass, field
from dataclasses import dataclass
import subprocess
import logging
logger = logging.getLogger()
@@ -774,5 +774,5 @@ def main():
return Result
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
-4
View File
@@ -1,13 +1,9 @@
#!/usr/bin/python3
from enum import Flag
import json
import os
import struct
import sys
import glob
from threading import Thread
import subprocess
import time
import multiprocessing
from shutil import which
+1 -1
View File
@@ -76,6 +76,6 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+2 -1
View File
@@ -1,7 +1,6 @@
#!/usr/bin/python3
import re
import sys
import subprocess
try:
from packaging.version import Version as version_check
except:
@@ -81,6 +80,8 @@ BigCoreIDs = {
[ ["apple-a13", "0.0"], # If we aren't on 12.0+
["apple-a14", "12.0"], # Only exists in 12.0+
],
# QEmu HVF 10.2+
tuple([0x61, 0]): "apple-a13", # Can't determine variant, choose lowest.
}
LittleCoreIDs = {
-1
View File
@@ -1,7 +1,6 @@
#!/bin/env python3
import sys
import fileinput
import re
# Handles the following formats:
+2 -2
View File
@@ -4,8 +4,8 @@ import os
import sys
import subprocess
# Check if FEX indicates support for AVX
def DoesFEXSupportAVX(mode):
# Check if FEX indicates support for AVX
fex_interpreter_path = os.path.dirname(sys.argv[7]) + "/FEX"
args = list()
@@ -22,8 +22,8 @@ def DoesFEXSupportAVX(mode):
return 'avx' in flags and 'avx2' in flags
return False
# Check if the test itself requires AVX
def TestRequiresAVXSupport():
# Check if the test itself requires AVX
exe_path = sys.argv[len(sys.argv) - 1]
json_path = os.path.dirname(os.path.dirname(exe_path)) + '/requirements/' + os.path.basename(exe_path) + '.json'
-3
View File
@@ -1,6 +1,3 @@
from enum import Flag
import json
import struct
import sys
from json_config_parse import parse_json
-3
View File
@@ -1,6 +1,3 @@
from enum import Flag
import json
import struct
import sys
from json_config_parse import parse_json
-3
View File
@@ -3,9 +3,6 @@
# Save current directory
DIR=$(pwd)
# Get the absolute path to the Scripts directory
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Parse arguments
CHANGED_ONLY=false
TARGET_DIR=""
+1 -2
View File
@@ -1,7 +1,6 @@
#!/usr/bin/python3
import sys
import subprocess
import os.path
from os import path
from shutil import which
@@ -77,4 +76,4 @@ if (is_known_failure):
sys.exit(1)
else:
# Just return the result code if we don't have this test as a known failure
sys.exit(ResultCode);
sys.exit(ResultCode)
+3
View File
@@ -14,10 +14,12 @@
#include <span>
#include <utility>
#include <vector>
#include <unistd.h>
#include <FEXCore/fextl/functional.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/LogManager.h>
namespace fasio {
@@ -368,6 +370,7 @@ std::size_t read(AsyncReadStream& Stream, mutable_buffer Buffers, error& ec) {
auto BytesRead = Stream.read_some(Buffers, ec);
TotalBytesRead += BytesRead;
if (Buffers.FD) {
LOGMAN_THROW_A_FMT(**Buffers.FD != -1, "Receiver requested a file descriptor but none was sent");
(void)Buffers.consume_fd();
}
Buffers += BytesRead;
+4 -2
View File
@@ -160,8 +160,10 @@ private:
struct cmsghdr* cmsg = CMSG_FIRSTHDR(&msg);
if (Buffers.FD &&
(cmsg == nullptr || cmsg->cmsg_len != CMSG_LEN(sizeof(int)) || cmsg->cmsg_level != SOL_SOCKET || cmsg->cmsg_type != SCM_RIGHTS)) {
ec = error::invalid;
return 0;
// Not a failure since some data was read for the main message
**Buffers.FD = -1;
ec = error::success;
return BytesRead;
}
if (Buffers.FD) {
+34 -2
View File
@@ -79,8 +79,8 @@ static char* SaveLayerToJSON(char* JsonBuffer, const FEXCore::Config::Layer* Lay
}
if (std::holds_alternative<fextl::string>(it.second)) {
JsonBuffer = json_str(JsonBuffer, Name.data(), std::get<fextl::string>(it.second).c_str());
} else if (std::holds_alternative<FEXCore::Config::DefaultValues::Type::StringArrayType>(it.second)) {
for (auto& var : std::get<FEXCore::Config::DefaultValues::Type::StringArrayType>(it.second)) {
} else if (std::holds_alternative<FEXCore::Config::StringArrayType>(it.second)) {
for (auto& var : std::get<FEXCore::Config::StringArrayType>(it.second)) {
JsonBuffer = json_str(JsonBuffer, Name.data(), var.c_str());
}
} else {
@@ -542,6 +542,13 @@ const char* GetHomeDirectory() {
#endif
fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* DataOverride = getenv("FEX_APP_DATA_LOCATION");
if (PortableInfo.IsPortable && (Global || !DataOverride)) {
@@ -566,6 +573,13 @@ fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableI
}
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (PortableInfo.IsPortable && (Global || !ConfigOverride)) {
return fextl::fmt::format("{}/fex-emu/", PortableInfo.InterpreterPath);
@@ -602,6 +616,24 @@ fextl::string GetConfigDirectory(bool Global, const PortableInformation& Portabl
return ConfigDir;
}
fextl::string GetCacheDirectory() {
const char* CacheOverride = getenv("FEX_APP_CACHE_LOCATION");
if (CacheOverride) {
return CacheOverride;
}
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_SHADER_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* HomeDir = GetHomeDirectory();
const char* CacheXDG = getenv("XDG_CACHE_HOME");
return (CacheXDG ? fextl::string {CacheXDG} : (fextl::string {HomeDir} + "/.cache")) + "/fex-emu/";
}
fextl::string GetConfigFileLocation(bool Global, const PortableInformation& PortableInfo) {
return GetConfigDirectory(Global, PortableInfo) + "Config.json";
}
+1
View File
@@ -62,6 +62,7 @@ const char* GetHomeDirectory();
fextl::string GetDataDirectory(const PortableInformation& PortableInfo);
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo);
fextl::string GetConfigFileLocation(bool Global, const PortableInformation& PortableInfo);
fextl::string GetCacheDirectory();
void InitializeConfigs(const PortableInformation& PortableInfo);
+166 -104
View File
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: MIT
#include "Common/AsyncNet.h"
#include "Common/Config.h"
#include "FDUtils.h"
#include "Common/FEXServerClient.h"
#include <FEXCore/Utils/CompilerDefs.h>
@@ -59,8 +60,8 @@ int RequestPIDFDPacket(int ServerSocket, PacketType Type) {
fasio::mutable_buffer ResBuffer {std::as_writable_bytes(std::span {&Res, 1})};
int NewFD = -1;
ResBuffer.FD = &NewFD;
read(Socket, ResBuffer, ec);
if (ec != fasio::error::success || Res.Header.Type != PacketType::TYPE_SUCCESS) {
auto BytesRead = Socket.read_some(ResBuffer, ec);
if (ec != fasio::error::success || BytesRead != sizeof(Res) || Res.Header.Type != PacketType::TYPE_SUCCESS) {
return -1;
}
@@ -137,15 +138,21 @@ fextl::string GetServerSocketName() {
}
fextl::string GetServerSocketPath() {
fextl::string name {};
#ifndef FEX_STEAM_SUPPORT
FEX_CONFIG_OPT(ServerSocketPath, SERVERSOCKETPATH);
auto name = ServerSocketPath();
name = ServerSocketPath();
if (name.starts_with("/")) {
return name;
}
auto Folder = GetTempFolder();
#else
// Under Steam the FEXServer's socket is a game-specific directory.
auto Folder = GetServerLockFolder();
#endif
if (name.empty()) {
return fextl::fmt::format("{}/{}.FEXServer.Socket", Folder, ::getuid());
@@ -159,25 +166,31 @@ int GetServerFD() {
}
int ConnectToServer(ConnectionOption ConnectionOption) {
auto ServerSocketName = GetServerSocketName();
int SocketFD {-1};
size_t SizeOfAddr {};
struct sockaddr_un addr {};
size_t SizeOfSocketString {};
// Create the initial unix socket
int SocketFD = socket(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0);
SocketFD = socket(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0);
if (SocketFD == -1) {
LogMan::Msg::EFmt("Couldn't open AF_UNIX socket {}", errno);
return -1;
}
// Steam doesn't get to connect to global sockets.
#ifndef FEX_STEAM_SUPPORT
auto ServerSocketName = GetServerSocketName();
// AF_UNIX has a special feature for named socket paths.
// If the name of the socket begins with `\0` then it is an "abstract" socket address.
// The entirety of the name is used as a path to a socket that doesn't have any filesystem backing.
struct sockaddr_un addr {};
addr.sun_family = AF_UNIX;
size_t SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
addr.sun_path[0] = 0; // Abstract AF_UNIX sockets start with \0
strncpy(addr.sun_path + 1, ServerSocketName.data(), SizeOfSocketString);
// Include final null character.
size_t SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
if (connect(SocketFD, reinterpret_cast<struct sockaddr*>(&addr), SizeOfAddr) == -1) {
if (ConnectionOption == ConnectionOption::Default || errno != ECONNREFUSED) {
@@ -186,11 +199,13 @@ int ConnectToServer(ConnectionOption ConnectionOption) {
} else {
return SocketFD;
}
#endif
// Try again with a path-based socket, since abstract sockets will fail if we have been
// placed in a new netns as part of a sandbox.
auto ServerSocketPath = GetServerSocketPath();
addr.sun_family = AF_UNIX;
SizeOfSocketString = std::min(ServerSocketPath.size(), sizeof(addr.sun_path) - 1);
strncpy(addr.sun_path, ServerSocketPath.data(), SizeOfSocketString);
SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
@@ -224,110 +239,125 @@ bool SetupClient(std::string_view InterpreterPath) {
return true;
}
int ConnectToAndStartServer(std::string_view InterpreterPath) {
int ServerFD = ConnectToServer(ConnectionOption::NoPrintConnectionError);
if (ServerFD == -1) {
// Couldn't connect to the server. Start one
int StartServer(std::string_view InterpreterPath, int watch_fd) {
int LocalServerFD {-1};
// Couldn't connect to the server. Start one
// Open some pipes for letting us know when the server is ready
int fds[2] {};
if (pipe2(fds, 0) != 0) {
LogMan::Msg::EFmt("Couldn't open pipe");
// Open some pipes for letting us know when the server is ready
int fds[2] {};
if (pipe2(fds, 0) != 0) {
LogMan::Msg::EFmt("Couldn't open pipe");
return -1;
}
// Extract directory from InterpreterPath
fextl::string InterpreterDir {InterpreterPath};
size_t LastSlash = InterpreterDir.rfind('/');
if (LastSlash != fextl::string::npos) {
InterpreterDir = InterpreterDir.substr(0, LastSlash);
}
fextl::string FEXServerPath = fextl::fmt::format("{}/FEXServer", InterpreterDir);
// Check if a local FEXServer next to FEX exists
// If it does then it takes priority over the installed one
if (!FHU::Filesystem::Exists(FEXServerPath)) {
FEXServerPath = "FEXServer";
}
// Set-up our SIGCHLD handler to ignore the signal.
// This is early in the initialization stage so no handlers have been installed.
//
// We want to ignore the signal so that if FEXServer starts in daemon mode, it
// doesn't leave a zombie process around waiting for something to get the result.
struct sigaction action {};
action.sa_handler = SIG_IGN;
sigaction(SIGCHLD, &action, &action);
pid_t pid = fork();
if (pid == 0) {
// Child
close(fds[0]); // Close read end of pipe
const char* argv[6];
auto pipe_string = fextl::fmt::format("{}", fds[1]);
auto watch_fd_string = fextl::fmt::format("{}", watch_fd);
size_t arg_count {};
argv[arg_count++] = FEXServerPath.c_str();
argv[arg_count++] = "--wait_pipe";
argv[arg_count++] = pipe_string.c_str();
if (watch_fd != -1) {
argv[arg_count++] = "--watch_fd";
argv[arg_count++] = watch_fd_string.c_str();
}
argv[arg_count++] = nullptr;
if (execvp(argv[0], (char* const*)argv) == -1) {
// Let the parent know that we couldn't execute for some reason
uint64_t error {1};
write(fds[1], &error, sizeof(error));
// Give a hopefully helpful error message for users
LogMan::Msg::EFmt("Couldn't execute: {}", argv[0]);
LogMan::Msg::EFmt("This means the squashFS rootfs won't be mounted.");
LogMan::Msg::EFmt("Expect errors!");
// Destroy this fork
exit(1);
}
FEX_UNREACHABLE;
} else {
// Parent
// Wait for the child to exit so we can check if it is mounted or not
close(fds[1]); // Close write end of the pipe
// Wait for a message from FEXServer
pollfd PollFD;
PollFD.fd = fds[0];
PollFD.events = POLLIN | POLLOUT | POLLRDHUP | POLLERR | POLLHUP | POLLNVAL;
// Wait for a result on the pipe that isn't EINTR
while (poll(&PollFD, 1, -1) == -1 && errno == EINTR)
;
// Check if child signaled an error
uint64_t error = 0;
ssize_t bytes_read = read(fds[0], &error, sizeof(error));
close(fds[0]);
if (bytes_read > 0 && error != 0) {
return -1;
}
// Extract directory from InterpreterPath
fextl::string InterpreterDir {InterpreterPath};
size_t LastSlash = InterpreterDir.rfind('/');
if (LastSlash != fextl::string::npos) {
InterpreterDir = InterpreterDir.substr(0, LastSlash);
for (size_t i = 0; i < 5; ++i) {
LocalServerFD = ConnectToServer(ConnectionOption::Default);
if (LocalServerFD != -1) {
break;
}
std::this_thread::sleep_for(std::chrono::seconds(1));
}
fextl::string FEXServerPath = fextl::fmt::format("{}/FEXServer", InterpreterDir);
// Check if a local FEXServer next to FEX exists
// If it does then it takes priority over the installed one
if (!FHU::Filesystem::Exists(FEXServerPath)) {
FEXServerPath = "FEXServer";
if (LocalServerFD == -1) {
// Still couldn't connect to the socket.
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket after launching the process");
}
// Set-up our SIGCHLD handler to ignore the signal.
// This is early in the initialization stage so no handlers have been installed.
//
// We want to ignore the signal so that if FEXServer starts in daemon mode, it
// doesn't leave a zombie process around waiting for something to get the result.
struct sigaction action {};
action.sa_handler = SIG_IGN;
sigaction(SIGCHLD, &action, &action);
pid_t pid = fork();
if (pid == 0) {
// Child
close(fds[0]); // Close read end of pipe
const char* argv[4];
auto pipe_string = fextl::fmt::format("{}", fds[1]);
argv[0] = FEXServerPath.c_str();
argv[1] = "--wait_pipe";
argv[2] = pipe_string.c_str();
argv[3] = nullptr;
if (execvp(argv[0], (char* const*)argv) == -1) {
// Let the parent know that we couldn't execute for some reason
uint64_t error {1};
write(fds[1], &error, sizeof(error));
// Give a hopefully helpful error message for users
LogMan::Msg::EFmt("Couldn't execute: {}", argv[0]);
LogMan::Msg::EFmt("This means the squashFS rootfs won't be mounted.");
LogMan::Msg::EFmt("Expect errors!");
// Destroy this fork
exit(1);
}
FEX_UNREACHABLE;
} else {
// Parent
// Wait for the child to exit so we can check if it is mounted or not
close(fds[1]); // Close write end of the pipe
// Wait for a message from FEXServer
pollfd PollFD;
PollFD.fd = fds[0];
PollFD.events = POLLIN | POLLOUT | POLLRDHUP | POLLERR | POLLHUP | POLLNVAL;
// Wait for a result on the pipe that isn't EINTR
while (poll(&PollFD, 1, -1) == -1 && errno == EINTR)
;
// Check if child signaled an error
uint64_t error = 0;
ssize_t bytes_read = read(fds[0], &error, sizeof(error));
close(fds[0]);
if (bytes_read > 0 && error != 0) {
return -1;
}
for (size_t i = 0; i < 5; ++i) {
ServerFD = ConnectToServer(ConnectionOption::Default);
if (ServerFD != -1) {
break;
}
std::this_thread::sleep_for(std::chrono::seconds(1));
}
if (ServerFD == -1) {
// Still couldn't connect to the socket.
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket {} after launching the process", GetServerSocketName());
}
}
// Restore the original SIGCHLD handler if it existed.
sigaction(SIGCHLD, &action, nullptr);
}
return ServerFD;
// Restore the original SIGCHLD handler if it existed.
sigaction(SIGCHLD, &action, nullptr);
return LocalServerFD;
}
int ConnectToAndStartServer(std::string_view InterpreterPath) {
int LocalServerFD = ConnectToServer(ConnectionOption::NoPrintConnectionError);
if (LocalServerFD == -1) {
LocalServerFD = StartServer(InterpreterPath);
}
return LocalServerFD;
}
/**
@@ -375,6 +405,38 @@ int RequestPIDFD(int ServerSocket) {
return RequestPIDFDPacket(ServerSocket, PacketType::TYPE_GET_PID_FD);
}
int RequestCodeMapFD(int ServerSocket, int ProgramFD, bool HasMultiblock) {
fasio::tcp_socket Socket {ServerSocket};
FEXServerRequestPacket Req {
.Header {
.Type = HasMultiblock ? PacketType::TYPE_QUERY_CODE_MAP : PacketType::TYPE_QUERY_CODE_MAP_NO_MULTIBLOCK,
},
};
// Send request
fasio::error ec;
{
fasio::mutable_buffer WriteBuffer {std::as_writable_bytes(std::span {&Req, 1})};
WriteBuffer.FD = &ProgramFD;
write(Socket, WriteBuffer, ec);
if (ec != fasio::error::success) {
return -1;
}
}
// Wait for success response and log FD
FEXServerResultPacket Res {};
fasio::mutable_buffer ResBuffer {std::as_writable_bytes(std::span {&Res, 1})};
int NewFD = -1;
ResBuffer.FD = &NewFD;
read(Socket, ResBuffer, ec);
if (ec != fasio::error::success || Res.Header.Type != PacketType::TYPE_SUCCESS) {
return -1;
}
return NewFD;
}
/** @} */
/**
+19
View File
@@ -19,6 +19,8 @@ enum class PacketType {
TYPE_GET_LOG_FD,
TYPE_GET_ROOTFS_PATH,
TYPE_GET_PID_FD,
TYPE_QUERY_CODE_MAP,
TYPE_QUERY_CODE_MAP_NO_MULTIBLOCK,
// Result only
TYPE_SUCCESS,
@@ -65,6 +67,13 @@ int GetServerFD();
bool SetupClient(std::string_view InterpreterPath);
/**
* @brief Start a FEXServer instance if possible
*
* @return socket FD for communicating with server
*/
int StartServer(std::string_view InterpreterPath, int watch_fd = -1);
/**
* @brief Connect to and start a FEXServer instance if required
*
@@ -113,6 +122,16 @@ fextl::string RequestRootFSPath(int ServerSocket);
*/
int RequestPIDFD(int ServerSocket);
/**
* @brief Request FEXServer to create a new code map for disk cache population
*
* @param ServerSocket - Socket to the server
* @param ProgramFD - FD for program binary
*
* @return FD to write code map to
*/
int RequestCodeMapFD(int ServerSocket, int ProgramFD, bool HasMultiblock);
/** @} */
/**
+9 -5
View File
@@ -2,9 +2,9 @@
#pragma once
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <optional>
#include <span>
#include <elf.h>
@@ -15,12 +15,16 @@ namespace FEXCore {
* Infers the base virtual address from a file mapping (as described by parameters to a single
* call to mmap()).
*
* Usually the base address can uniquely be inferred, but in edge cases multiple possible
* candidates are returned.
*
* The file offset of any given mapping need not match its virtual address offset from the base
* mapping (file offset = 0). Instead, this function searches the corresponding ELF program headers
* for an entry that generated the given file mapping.
*/
inline std::optional<uint64_t>
inline fextl::vector<uint64_t>
InferMappingBaseAddress(std::span<const Elf64_Phdr> ProgramHeaders, uint64_t Addr, uint64_t Size, uint64_t FileOffset, int AccessFlags) {
fextl::vector<uint64_t> Ret;
for (auto& phdr : ProgramHeaders) {
if (phdr.p_type != PT_LOAD) {
// Skip headers that don't trigger memory mappings
@@ -36,11 +40,11 @@ InferMappingBaseAddress(std::span<const Elf64_Phdr> ProgramHeaders, uint64_t Add
if (FileOffset >= SegmentStartOffset && FileOffset < SegmentStartOffset + phdr.p_filesz &&
(FileOffset & Utils::FEX_PAGE_MASK) == (phdr.p_offset & Utils::FEX_PAGE_MASK)) {
// Compute VA offset relative to the base mapping
return Addr - (phdr.p_vaddr - (phdr.p_offset & 0xfff)) + (ProgramHeaders[0].p_vaddr - (ProgramHeaders[0].p_offset & 0xfff)) -
(FileOffset - SegmentStartOffset);
Ret.push_back(Addr - (phdr.p_vaddr - (phdr.p_offset & 0xfff)) + (ProgramHeaders[0].p_vaddr - (ProgramHeaders[0].p_offset & 0xfff)) -
(FileOffset - SegmentStartOffset));
}
}
return std::nullopt;
return Ret;
}
} // namespace FEXCore
+193 -69
View File
@@ -7,6 +7,9 @@
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/StringUtils.h>
#include <range/v3/view/split.hpp>
#include <range/v3/view/transform.hpp>
#ifdef _M_X86_64
#include "Common/X86Features.h"
#endif
@@ -35,6 +38,22 @@ void FillMIDRInformationViaLinux(FEXCore::HostFeatures* Features) {
#endif
}
#if defined(_M_ARM_64) && !defined(VIXL_SIMULATOR)
__attribute__((naked)) static uint64_t ReadSVEVectorLengthInBits() {
///< Can't use rdvl instruction directly because compilers will complain that sve/sme is required.
__asm(R"(
.word 0x04bf5100 // rdvl x0, #8
ret;
)");
}
#else
[[maybe_unused]]
static int ReadSVEVectorLengthInBits() {
// Return unsupported
return 0;
}
#endif
#ifdef _M_ARM_64
#define GetSysReg(name, reg) \
static uint64_t Get_##name() { \
@@ -53,6 +72,7 @@ GetSysReg(MMFR2_EL1, ID_AA64MMFR2_EL1);
GetSysReg(ZFR0_EL1, s3_0_c0_c4_4); // Can't request by name
GetSysReg(MMFR1_EL1, ID_AA64MMFR1_EL1);
GetSysReg(ISAR2_EL1, ID_AA64ISAR2_EL1);
GetSysReg(DCZID_EL0, DCZID_EL0);
class CPUFeaturesFromID final : public FEX::CPUFeatures {
public:
@@ -66,12 +86,17 @@ public:
MMFR2.SetReg(Get_MMFR2_EL1());
MMFR1.SetReg(Get_MMFR1_EL1());
ISAR2.SetReg(Get_ISAR2_EL1());
DCZID.SetReg(Get_DCZID_EL0());
if (PFR0.SupportsSVE()) {
// Can only query if SVE is supported.
ZFR0.SetReg(Get_ZFR0_EL1());
}
FillFeatureFlags();
if (Supports(CPUFeatures::Feature::SVE2)) {
SVEVL.SetReg(ReadSVEVectorLengthInBits());
}
}
};
@@ -80,6 +105,78 @@ FEX::CPUFeatures GetCPUFeaturesFromIDRegisters() {
}
#endif
class CPUFeaturesFromConfig final : public FEX::CPUFeatures {
public:
CPUFeaturesFromConfig(std::string_view Config) {
auto to_string_view = [](auto rng) {
return std::string_view(&*rng.begin(), ranges::distance(rng));
};
for (auto Option : ranges::views::split(Config, ',') | ranges::views::transform(to_string_view)) {
auto OptionData = ranges::views::split(Option, '=') | ranges::views::transform(to_string_view);
auto OptionDataBegin = ranges::begin(OptionData);
auto OptionDataEnd = ranges::end(OptionData);
if (OptionDataBegin == OptionDataEnd) {
continue;
}
auto Key = *OptionDataBegin;
if (Key.empty()) {
continue;
}
++OptionDataBegin;
if (OptionDataBegin == OptionDataEnd) {
continue;
}
auto Value = *OptionDataBegin;
uint64_t ValueHex {};
char* str_end {};
ValueHex = std::strtoull(Value.data(), &str_end, 16);
if (str_end == Value.data()) {
LogMan::Msg::EFmt("Couldn't parse '{}={}'\n", Key, Value);
continue;
}
if (Key == "isar0") {
ISAR0.SetReg(ValueHex);
} else if (Key == "isar1") {
ISAR1.SetReg(ValueHex);
} else if (Key == "isar2") {
ISAR2.SetReg(ValueHex);
} else if (Key == "pfr0") {
PFR0.SetReg(ValueHex);
} else if (Key == "pfr1") {
PFR1.SetReg(ValueHex);
} else if (Key == "midr") {
MIDR.SetReg(ValueHex);
} else if (Key == "mmfr0") {
MMFR0.SetReg(ValueHex);
} else if (Key == "mmfr1") {
MMFR1.SetReg(ValueHex);
} else if (Key == "mmfr2") {
MMFR2.SetReg(ValueHex);
} else if (Key == "zfr0") {
ZFR0.SetReg(ValueHex);
} else if (Key == "dczid") {
DCZID.SetReg(ValueHex);
} else if (Key == "svevl") {
SVEVL.SetReg(ValueHex);
} else {
LogMan::Msg::EFmt("Unknown Key: {}", Key);
}
}
FillFeatureFlags();
}
};
FEX::CPUFeatures GetCPUFeaturesFromConfig(std::string_view Config) {
return CPUFeaturesFromConfig {Config};
}
class CPUFeaturesAll final : public FEX::CPUFeatures {
public:
CPUFeaturesAll() {
@@ -87,6 +184,9 @@ public:
for (uint32_t i = 0; i < FEXCore::ToUnderlying(FEX::CPUFeatures::Feature::MAX); ++i) {
SetFeature(FEX::CPUFeatures::Feature {i});
}
// Report unsupported for DCZVA
DCZID.SetReg(0b1'0000);
}
};
@@ -346,21 +446,7 @@ void FEX::CPUFeatures::FillFeatureFlags() {
}
}
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
[[maybe_unused]] constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
[[maybe_unused]] constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
[[maybe_unused]]
static uint32_t GetDCZID() {
uint64_t Result {};
__asm("mrs %[Res], DCZID_EL0" : [Res] "=r"(Result));
return Result;
}
static uint32_t GetFPCR() {
uint64_t Result {};
__asm("mrs %[Res], FPCR" : [Res] "=r"(Result));
@@ -371,27 +457,6 @@ static void SetFPCR(uint64_t Value) {
__asm("msr FPCR, %[Value]" ::[Value] "r"(Value));
}
#ifndef VIXL_SIMULATOR
__attribute__((naked)) static uint64_t ReadSVEVectorLengthInBits() {
///< Can't use rdvl instruction directly because compilers will complain that sve/sme is required.
__asm(R"(
.word 0x04bf5100 // rdvl x0, #8
ret;
)");
}
#endif
#else
[[maybe_unused]]
static uint32_t GetDCZID() {
// Return unsupported
return DCZID_DZP_MASK;
}
[[maybe_unused]]
static int ReadSVEVectorLengthInBits() {
// Return unsupported
return 0;
}
#endif
static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEWidth) {
@@ -460,9 +525,76 @@ static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEW
Features->SupportsSVE256 = ForceSVEWidth && ForceSVEWidth >= 256;
}
FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool SupportsCacheMaintenanceOps, uint64_t CTR, uint64_t MIDR) {
FEXCore::HostFeatures HostFeatures;
static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
constexpr uint32_t Implementer_ARM = 0x41;
constexpr uint32_t PartNum_V2 = 0xd4f;
constexpr uint32_t PartNum_V3 = 0xd84;
constexpr uint32_t PartNum_V3AE = 0xd83;
constexpr uint32_t PartNum_X3 = 0xd4e;
constexpr uint32_t PartNum_X4 = 0xd82;
constexpr uint32_t PartNum_X925 = 0xd85;
constexpr uint32_t PartNum_C1Ultra = 0xd8c;
constexpr uint32_t PartNum_C1Premium = 0xd90;
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
auto GetMIDRImplementer = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 24) & 0xFF;
};
auto GetMIDRPartNum = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 4) & 0xFFF;
};
const uint32_t MIDR_Implementer = GetMIDRImplementer(MIDR);
const uint32_t MIDR_PartNum = GetMIDRPartNum(MIDR);
#ifdef _M_ARM_64
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
// - The RNDRRS register will never read a random number. (Always return failure)
// This is contrary to x86 RNG behaviour where it allows spurious failure with RDSEED, but guarantees eventual success.
// This manifested itself on Linux when an x86 processor failed to guarantee forward progress and boot of services would infinite
// loop. Just disable this extension if this CPU is detected.
HostFeatures->SupportsRAND = false;
}
#endif
// The LDAPUR instruction suffers from significant performance issues on many ARM implementations. This is
// listed in the official Cortex errata list as follows:
//
// 3877900
// LDAPUR, LDAPURB, LDAPURH instructions have stricter memory ordering than required
//
// LDAPUR instructions execute with full Load-Acquire ordering instead of the relaxed ordering described
// in the LDAPUR pseudocode. This might cause significant performance degradation in workloads that do
// not require this stricter memory ordering. Note that this erratum only affects the unscaled versions of
// LDAPUR (LDAPUR, LDAPURB, LDAPURH), and not LDAPR (LDAPR, LDAPRB, LDAPRH).
//
// The list of cores to disable its use on was taken from the following LLVM PR that accomplishes the same
// thing: https://github.com/llvm/llvm-project/pull/124274
for (uint32_t CoreIndex = 0; CoreIndex < HostFeatures->CPUMIDRs.size(); CoreIndex++) {
const uint32_t CoreMIDR = HostFeatures->CPUMIDRs[CoreIndex];
const uint32_t Core_MIDR_Implementer = GetMIDRImplementer(CoreMIDR);
const uint32_t Core_MIDR_PartNum = GetMIDRPartNum(CoreMIDR);
bool IgnoreLRCPC2 = (Core_MIDR_Implementer == Implementer_ARM) &&
((Core_MIDR_PartNum == PartNum_V2) || (Core_MIDR_PartNum == PartNum_V3) || (Core_MIDR_PartNum == PartNum_X3) ||
(Core_MIDR_PartNum == PartNum_X4) || (Core_MIDR_PartNum == PartNum_X925) || (Core_MIDR_PartNum == PartNum_V3AE) ||
(Core_MIDR_PartNum == PartNum_C1Ultra) || (Core_MIDR_PartNum == PartNum_C1Premium));
if (IgnoreLRCPC2) {
HostFeatures->SupportsTSOImm9 = false;
break;
}
}
}
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
uint64_t MIDR) {
FEX_CONFIG_OPT(ForceSVEWidth, FORCESVEWIDTH);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
@@ -495,7 +627,7 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
HostFeatures.SupportsSVE256 = ForceSVEWidth() ? ForceSVEWidth() >= 256 : true;
#else
HostFeatures.SupportsSVE128 = Features.Supports(CPUFeatures::Feature::SVE2);
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && ReadSVEVectorLengthInBits() >= 256;
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && Features.GetSVEVectorLengthInBits() >= 256;
#endif
HostFeatures.SupportsAVX = true;
@@ -532,23 +664,6 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
// Set FPCR back to original just in case anything changed
SetFPCR(OriginalFPCR);
if (HostFeatures.SupportsRAND) {
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
const uint32_t MIDR_Implementer = (MIDR >> 24) & 0xFF;
const uint32_t MIDR_PartNum = (MIDR >> 4) & 0xFFF;
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
// - The RNDRRS register will never read a random number. (Always return failure)
// This is contrary to x86 RNG behaviour where it allows spurious failure with RDSEED, but guarantees eventual success.
// This manifested itself on Linux when an x86 processor failed to guarantee forward progress and boot of services would infinite
// loop. Just disable this extension if this CPU is detected.
HostFeatures.SupportsRAND = false;
}
}
#endif
#ifdef VIXL_SIMULATOR
@@ -560,16 +675,17 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
HostFeatures.SupportsSHA = true;
HostFeatures.SupportsPMULL_128Bit = true;
HostFeatures.SupportsAES256 = true;
// Simulator doesn't support these
HostFeatures.SupportsRPRES = false;
HostFeatures.SupportsAFP = false;
#else
// Check if we can support cacheline clears
uint32_t DCZID = GetDCZID();
if ((DCZID & DCZID_DZP_MASK) == 0) {
uint32_t DCZID_Log2 = DCZID & DCZID_BS_MASK;
uint32_t DCZID_Bytes = (1 << DCZID_Log2) * sizeof(uint32_t);
if (Features.GetDCZID().SupportsDCZVA()) {
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
constexpr static uint64_t CACHELINE_SIZE = 64;
HostFeatures.SupportsCLZERO = DCZID_Bytes == CACHELINE_SIZE;
HostFeatures.SupportsCLZERO = Features.GetDCZID().BlockSizeInBytes() == CACHELINE_SIZE;
}
#endif
@@ -598,21 +714,28 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
#endif
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
HandleErrata(&HostFeatures, MIDR);
OverrideFeatures(&HostFeatures, ForceSVEWidth());
return HostFeatures;
}
FEXCore::HostFeatures FetchHostFeatures() {
#ifdef _M_X86_64
CPUFeatures Features = CPUFeaturesAll {};
FEX_CONFIG_OPT(CPUFeatureRegisters, CPUFEATUREREGISTERS);
// Vixl simulator doesn't support AFP.
Features.RemoveFeature(CPUFeatures::Feature::AFP);
// Vixl simulator doesn't support RPRES.
Features.RemoveFeature(CPUFeatures::Feature::RPRES);
CPUFeatures Features {};
if (!CPUFeatureRegisters().empty()) {
Features = GetCPUFeaturesFromConfig(CPUFeatureRegisters());
} else {
#ifdef _M_X86_64
Features = CPUFeaturesAll {};
// Vixl simulator doesn't support AFP.
Features.RemoveFeature(CPUFeatures::Feature::AFP);
// Vixl simulator doesn't support RPRES.
Features.RemoveFeature(CPUFeatures::Feature::RPRES);
#else
CPUFeatures Features = GetCPUFeaturesFromIDRegisters();
Features = GetCPUFeaturesFromIDRegisters();
#endif
}
uint64_t CTR = 0;
uint64_t MIDR = 0;
@@ -623,8 +746,9 @@ FEXCore::HostFeatures FetchHostFeatures() {
__asm volatile("mrs %[midr], midr_el1" : [midr] "=r"(MIDR));
#endif
auto HostFeatures = FetchHostFeatures(Features, true, CTR, MIDR);
FEXCore::HostFeatures HostFeatures = {};
FillMIDRInformationViaLinux(&HostFeatures);
FetchHostFeatures(Features, HostFeatures, true, CTR, MIDR);
HostFeatures.SupportsCPUIndexInTPIDRRO = false;
return HostFeatures;
+69 -30
View File
@@ -8,6 +8,23 @@
namespace FEX {
class CPUFeatures {
public:
class FeatureReg {
public:
void SetReg(uint64_t _Reg) {
Reg = _Reg;
}
uint64_t Get() const {
return Reg;
}
protected:
// All feature flag fields are 4-bits.
uint64_t GetField(uint64_t Offset) const {
return (Reg >> Offset) & 0b1111;
}
uint64_t Reg {};
};
enum class Feature : uint32_t {
// ISAR0
AES,
@@ -100,23 +117,25 @@ public:
MAX,
};
static_assert(FEXCore::ToUnderlying(Feature::MAX) < 128);
static_assert((FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) == 1);
class DCZIDReg final : public FeatureReg {
public:
bool SupportsDCZVA() const {
return (Reg & DCZID_DZP_MASK) == 0;
}
bool Supports(Feature feat) const {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
return (FeatureBits[DWordSelect] >> BitSelect) & 1;
}
uint32_t BlockSizeInBytes() const {
uint32_t DCZID_Log2 = Reg & DCZID_BS_MASK;
return (1 << DCZID_Log2) * sizeof(uint32_t);
}
void RemoveFeature(Feature feat) {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
FeatureBits[DWordSelect] &= ~(1ULL << BitSelect);
}
protected:
void FillFeatureFlags();
private:
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
[[maybe_unused]] constexpr static uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
[[maybe_unused]] constexpr static uint32_t DCZID_BS_MASK = 0b0'1111;
};
// This list is informed by Linux kernel's `Documentation/arch/arm64/cpu-feature-registers.rst`
enum class FeatureRegType {
@@ -132,24 +151,11 @@ protected:
ISAR2_EL1,
};
class FeatureReg {
public:
void SetReg(uint64_t _Reg) {
Reg = _Reg;
}
protected:
// All feature flag fields are 4-bits.
uint64_t GetField(uint64_t Offset) const {
return (Reg >> Offset) & 0b1111;
}
uint64_t Reg {};
};
#define FIELD_FETCHER(feature, field, minimum_field) \
bool Supports##feature() const { \
return GetField(field) >= minimum_field; \
}
class ISAR0Reg final : public FeatureReg {
public:
FIELD_FETCHER(AES, AES, 0b0001);
@@ -601,8 +607,11 @@ protected:
ATS1A = 15 * 4,
};
};
class SVEVLReg final : public FeatureReg {};
#undef FIELD_FETCHER
ISAR0Reg ISAR0;
PFR0Reg PFR0;
PFR1Reg PFR1;
@@ -613,6 +622,34 @@ protected:
MMFR2Reg MMFR2;
MMFR1Reg MMFR1;
ISAR2Reg ISAR2;
DCZIDReg DCZID;
SVEVLReg SVEVL;
static_assert(FEXCore::ToUnderlying(Feature::MAX) < 128);
static_assert((FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) == 1);
bool Supports(Feature feat) const {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
return (FeatureBits[DWordSelect] >> BitSelect) & 1;
}
void RemoveFeature(Feature feat) {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
FeatureBits[DWordSelect] &= ~(1ULL << BitSelect);
}
const DCZIDReg& GetDCZID() const {
return DCZID;
}
uint64_t GetSVEVectorLengthInBits() const {
return SVEVL.Get();
}
protected:
void FillFeatureFlags();
uint64_t FeatureBits[(FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) + 1] {};
@@ -625,6 +662,8 @@ protected:
void FillMIDRInformationViaLinux(FEXCore::HostFeatures* Features);
FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool SupportsCacheMaintenanceOps, uint64_t CTR, uint64_t MIDR);
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
uint64_t MIDR);
FEXCore::HostFeatures FetchHostFeatures();
FEX::CPUFeatures GetCPUFeaturesFromIDRegisters();
} // namespace FEX
+33
View File
@@ -0,0 +1,33 @@
add_executable(FEXCompatTool
CompatTool.cpp)
target_link_libraries(FEXCompatTool
PRIVATE
FEXCore Common CommonTools JemallocLibs)
install(TARGETS FEXCompatTool
RUNTIME
DESTINATION /
COMPONENT Runtime
)
add_executable(FEXServerManager
ServerManager.cpp)
target_link_libraries(FEXServerManager
PRIVATE
FEXCore Common CommonTools JemallocLibs)
install(TARGETS FEXServerManager
RUNTIME
DESTINATION bin
COMPONENT Runtime
)
# Description json gets installed into root of depot
install(FILES emulator.json
DESTINATION /
COMPONENT Runtime)
install(FILES ConfigTemplate.json
DESTINATION /
COMPONENT Runtime)
+197
View File
@@ -0,0 +1,197 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: Bin|FEXCompatTool
desc: Used for launching games from Steam
$end_info$
*/
#include "PortabilityInfo.h"
#include "Common/Config.h"
#include "FEXCore/Utils/FileLoading.h"
#include "FEXCore/Utils/StringUtils.h"
#include "FEXHeaderUtils/Filesystem.h"
#include <stdlib.h>
#include <tiny-json.h>
fextl::string GenerateSteamConfigTemplate(const FEX::Config::PortableInformation& PortableInfo) {
const auto ConfigTemplatePath = PortableInfo.InterpreterPath + "ConfigTemplate.json";
if (!FHU::Filesystem::Exists(ConfigTemplatePath)) {
return {};
}
fextl::string Data;
if (!FEXCore::FileLoading::LoadFile(Data, ConfigTemplatePath)) {
return {};
}
// Try and find a mount point.
fextl::string MountPoint {};
const char* RuntimeDir = getenv("XDG_RUNTIME_DIR");
if (RuntimeDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", RuntimeDir);
} else {
const auto UserDirectory = fextl::fmt::format("/run/user/{}", geteuid());
if (FHU::Filesystem::Exists(UserDirectory)) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", UserDirectory);
} else {
const char* CacheDir = getenv("XDG_CACHE_HOME");
if (CacheDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", CacheDir);
} else {
// We tried really hard to find a mount path.
MountPoint = "~/.cache/fexrootfs/";
}
}
}
// Update the @FEX_COMPAT_TOOL@ config to point to the root of the depot.
FEXCore::StringUtils::ReplaceAllInPlace(Data, "@FEX_COMPAT_TOOL@", PortableInfo.InterpreterPath);
// TODO: This path is getting phased out.
FEXCore::StringUtils::ReplaceAllInPlace(Data, "@FEX_ROOTFS_PATH@", MountPoint);
// Save the json.
const auto ConfigPath = FEX::Config::GetConfigDirectory(false, PortableInfo);
const auto ConfigLocation = ConfigPath + "Config.json";
if (!FHU::Filesystem::CreateDirectories(ConfigPath)) {
return {};
}
auto File = FEXCore::File::File(ConfigLocation.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
if (!File.IsValid()) {
return {};
}
File.Write(Data.data(), Data.size());
return ConfigPath;
}
fextl::string GenerateSteamAppConfig(const FEX::Config::PortableInformation& PortableInfo) {
const auto user_config = getenv("FEX_APP_CONFIG");
if (user_config) {
// If user supplied config then don't use Steam config.
return {};
}
// Current supported Steam options.
struct SteamOptions {
bool TSO = true;
bool Multiblock = true;
bool Thunks_GL = false;
bool Thunks_Vulkan = false;
bool EnableLogging = false;
};
SteamOptions Options {};
// Game overrides.
const auto steam_fex_tso = getenv("STEAM_FEX_TSOENABLED");
if (steam_fex_tso) {
Options.TSO = std::strtoull(steam_fex_tso, nullptr, 0) != 0;
}
const auto steam_fex_multiblock = getenv("STEAM_FEX_MULTIBLOCK");
if (steam_fex_multiblock) {
Options.Multiblock = std::strtoull(steam_fex_multiblock, nullptr, 0) != 0;
}
const auto steam_fex_logging = getenv("STEAM_FEX_LOG");
if (steam_fex_logging) {
Options.EnableLogging = std::strtoull(steam_fex_logging, nullptr, 0) != 0;
}
// UI overrides.
const auto steam_fex_compat = getenv("STEAM_COMPAT_FEX_CONFIG");
if (steam_fex_compat) {
const auto steam_fex_compat_view = std::string_view(steam_fex_compat);
if (steam_fex_compat_view.find("TSOEnabled:1") != steam_fex_compat_view.npos) {
Options.TSO = true;
}
if (steam_fex_compat_view.find("Multiblock:1") != steam_fex_compat_view.npos) {
Options.Multiblock = true;
}
if (steam_fex_compat_view.find("ThunksDB_GL:1") != steam_fex_compat_view.npos) {
Options.Thunks_GL = true;
}
if (steam_fex_compat_view.find("ThunksDB_Vulkan:1") != steam_fex_compat_view.npos) {
Options.Thunks_Vulkan = true;
}
}
// Create the json.
char Buffer[4096];
char* Dest {};
Dest = json_objOpen(Buffer, nullptr);
{
Dest = json_objOpen(Dest, "Config");
Dest = json_str(Dest, "TSOEnabled", Options.TSO ? "1" : "0");
Dest = json_str(Dest, "Multiblock", Options.Multiblock ? "1" : "0");
Dest = json_str(Dest, "SilentLog", Options.EnableLogging ? "0" : "1");
if (Options.EnableLogging) {
Dest = json_str(Dest, "OutputLog", "server");
}
Dest = json_objClose(Dest);
}
{
Dest = json_objOpen(Dest, "ThunksDB");
Dest = json_str(Dest, "GL", Options.Thunks_GL ? "1" : "0");
Dest = json_str(Dest, "Vulkan", Options.Thunks_Vulkan ? "1" : "0");
Dest = json_objClose(Dest);
}
Dest = json_objClose(Dest);
json_end(Dest);
// Save the json.
const auto ConfigPath = FEX::Config::GetConfigDirectory(false, PortableInfo);
const auto ConfigLocation = ConfigPath + "app_config.json";
if (!FHU::Filesystem::CreateDirectories(ConfigPath)) {
return {};
}
auto File = FEXCore::File::File(ConfigLocation.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
if (!File.IsValid()) {
return {};
}
File.Write(Buffer, strlen(Buffer));
return ConfigLocation;
}
int main(int argc, const char** argv) {
const auto PortableInfo = FEX::ReadPortabilityInformation();
const auto TemplateConfigPath = GenerateSteamConfigTemplate(PortableInfo);
const auto AppConfigPath = GenerateSteamAppConfig(PortableInfo);
if (!TemplateConfigPath.empty()) {
setenv("FEX_APP_CONFIG_LOCATION", TemplateConfigPath.c_str(), true);
}
if (!AppConfigPath.empty()) {
setenv("FEX_APP_CONFIG", AppConfigPath.c_str(), true);
}
const auto FEXInterpreterPath = PortableInfo.InterpreterPath + "usr/bin/FEX";
// Due to no arguments for this application, just replace argv[0] and execve again.
argv[0] = FEXInterpreterPath.c_str();
execv(FEXInterpreterPath.c_str(), const_cast<char* const*>(argv));
// Save errno as it can change after calling `perror`.
const auto saved_errno = errno;
perror(argv[0]);
if (saved_errno == ENOENT) {
return 127;
}
return 126;
}
+9
View File
@@ -0,0 +1,9 @@
{
"Config": {
"X87ReducedPrecision": "1",
"RootFS": "@FEX_ROOTFS_PATH@/",
"ThunkHostLibs": "@FEX_COMPAT_TOOL@/usr/lib/aarch64-linux-gnu/fex-emu/HostThunks",
"ThunkGuestLibs": "@FEX_COMPAT_TOOL@/usr/share/fex-emu/GuestThunks",
"ProfileStats": "1"
}
}
Loaded 100 of 225 files, more files were not shown because too many files have changed in this diff. Show more