Compare commits

..
197 Commits
Author SHA1 Message Date
Ryan Houdek ba1b4744c5 Docs: Update for release FEX-2512 2025-12-05 16:11:40 -08:00
Ryan Houdek bd13c02451 Merge pull request #5092 from Sonicadvance1/12
SyscallsSMCTracking: Workaround assert in ELF mapping
2025-12-05 15:44:09 -08:00
Ryan Houdek 5902b175f9 SyscallsSMCTracking: Workaround assert in ELF mapping
The ELF tracking thing has an expectation that only portions of ELF
files that are described in the program headers will be mapped
executable. This doesn't hold true as programs will remap random
portions of ELF files as executable. In the case that this occurs, don't
assert out and instead print a warning.

This was discovered as Node.js remaps a portion of itself executable
that isn't described as such in the program headers. I also have a local
unittest that exposes the same problem. I had discovered same problem in
some other program with #5038.

Also fixes a bug where sometimes completely anonymously mapped
executable sneak in and cause a crash, which is kind of silly.
2025-12-05 14:50:58 -08:00
Ryan Houdek d0e47f9073 Merge pull request #5096 from neobrain/feature_fexofflinecompiler
CodeCache: Implement offline compiler for cache generation
2025-12-05 14:47:58 -08:00
Ryan Houdek bd7215d36f Merge pull request #5093 from Sonicadvance1/13
SteamRT4: Adds support for building the steam depot
2025-12-05 14:39:30 -08:00
Ryan Houdek f3f134f9de FEXServer: Add support for FEX Logging control
Always enables FEXServer log thread when built for Steam so that clients
can be controlled with `STEAM_FEX_LOG=1`. FEX logs will then always go
to FEXServer and those can get directed to wherever pressure-vessel
chooses.
2025-12-04 14:06:18 -08:00
Ryan Houdek cdbf5d57bd Steam: Adds FEXServerManager
This is fairly simple. Needs to be installed alongside FEXServer, so
that it can start it in a portable config.

- Starts a FEXServer
- Tells pressure-vessel when FEXServer is ready
- Keeps FEXServer alive with the `watch_fd` as long as the process lives
- Listens for pressure-vessel to be shutting down
- Exits once pressure-vessel exits, also letting FEXServer shutdown if
  no FEX instances are alive.
2025-12-04 14:06:17 -08:00
Ryan Houdek 1ae50bd670 FEXServerClient: Split out Connect and Start
Allow an optional watch_fd to be passed to FEXServer.
2025-12-04 14:02:43 -08:00
Ryan Houdek d28c9f9843 FEXServerClient: Disallow abstract named sockets under Steam
We don't want clients connecting to random sockets.
2025-12-04 14:02:43 -08:00
Ryan Houdek fe7d52aa78 github/steamrt4: Add artifacts 2025-12-04 14:02:43 -08:00
Ryan Houdek fc0907f8c1 Steam: Add FEXCompatTool 2025-12-04 14:02:43 -08:00
Ryan Houdek e57678d7a6 Config: Move Steam configs into config system 2025-12-04 14:02:43 -08:00
Ryan Houdek 45e594e806 Utils/StringUtils: Add in-place token replace helper 2025-12-04 14:02:43 -08:00
Ryan Houdek 87e7a0effa CMake: Don't install test thunk if tests aren't enabled 2025-12-04 14:02:43 -08:00
Ryan Houdek 4fd1a35b2a CMake: Disable some installs when building for Steam 2025-12-04 14:02:42 -08:00
Ryan Houdek c460cf0678 Merge pull request #5097 from wcampbell-nv/cmdline-map
Remap /proc/pid/cmdline with PR_SET_MM_MAP
2025-12-04 14:02:19 -08:00
Tony Wasserka 983802da61 CodeCache: Add offline compiler for generating caches 2025-12-04 19:16:51 +01:00
Tony Wasserka 49273e0d59 CodeCache: Add workaround for clang-15's broken std::piecewise_construct 2025-12-04 19:16:51 +01:00
Tony Wasserka 76b8459cdc CodeCache: Zero-initialize FileId in ExecutableFileInfo 2025-12-04 19:16:51 +01:00
Tony Wasserka e269eb6f65 ELFCodeLoader: Add helper interfaces 2025-12-04 19:16:51 +01:00
Tony Wasserka 7b4774f375 ELFCodeLoader: Add option to skip interpreter loading 2025-12-04 19:16:51 +01:00
Tony Wasserka 70e9a25112 Syscalls: Move m(un)map to a dedicated interface 2025-12-04 19:16:51 +01:00
Tony Wasserka 9fb83ea56a Windows/CRT: Implement lseek 2025-12-04 19:16:51 +01:00
LC 8bb3398376 Merge pull request #5101 from Sonicadvance1/16
SVE256: Fixes AVX scalar round with insert
2025-12-04 11:24:49 -05:00
Will Campbell d42fbb3d4d Use LoadFileToBuffer + cleanup 2025-12-04 07:54:45 -08:00
Tony Wasserka 53b2245dc1 Merge pull request #5098 from Sonicadvance1/14
docs: Update ProgrammingConcerns
2025-12-04 09:52:59 +00:00
Ryan Houdek db14975828 InstcountCI: Update 2025-12-04 01:45:32 -08:00
Ryan Houdek a5139d2710 unittests/ASM: Adds test case for AVX scalar round with insert bug 2025-12-04 01:44:12 -08:00
Ryan Houdek 9f584c8014 SVE256: Fixes AVX scalar round with insert
We were using the incorrect source registers on SVE256 implementation of
these instructions.

Fixes #5100
2025-12-04 01:43:07 -08:00
Tony Wasserka 6f1b98fb52 docs: Clean up ProgrammingConcerns 2025-12-04 10:18:26 +01:00
Ryan Houdek c258a90505 docs: Update ProgrammingConcerns
Disallow all APIs that touch `FILE`, they all allocate memory that we
don't control.
2025-12-03 20:32:33 -08:00
Will Campbell eb47ef43a7 Read from a fd rather than a FILE 2025-12-03 19:59:13 -08:00
Will Campbell 1876d6b923 Remap cmdline 2025-12-03 18:15:33 -08:00
Ryan Houdek 90c8fcf393 Merge pull request #5094 from pmatos/fix/issue5084
Set current code block in x87 pass
2025-12-02 14:43:54 -08:00
Ryan Houdek 6af90575e9 Merge pull request #5095 from neobrain/feature_serialize_relocations
JIT: Add support for serializing relocations
2025-12-02 14:43:42 -08:00
Tony Wasserka 994613260c CodeCache: Support reverse application of relocations
This allows code to be serialized consistently across runs.
2025-12-02 22:27:53 +01:00
Tony Wasserka 1430fa8220 CodeCache: Move ApplyCodeRelocations 2025-12-02 18:38:59 +01:00
Tony Wasserka 33e06058c6 JIT: Move ApplyRelocations to CodeCache 2025-12-02 18:38:59 +01:00
Tony Wasserka b032d1e1f7 JIT: Make relocations relative to guest base before serialization
This ensures consistency of generated code caches across multiple runs.
2025-12-02 18:38:59 +01:00
Tony Wasserka d6b43b1fe6 Dispatcher: Add public interface to query ExitFunctionLinkerAddress 2025-12-02 17:56:56 +01:00
Paulo Matos fc771c8683 asm_tests: Set current code block in x87 pass 2025-12-02 15:33:08 +01:00
Paulo Matos f91ac09f87 Set current code block in x87 pass
This resets the constant pool in IREmit used by SelectAddressMode().

Fixes #5084.
2025-12-02 15:33:08 +01:00
Ryan Houdek e4fa399412 Merge pull request #5091 from neobrain/feature_better_relocations
JIT: Prepare FEX relocations for code caching
2025-12-01 13:21:29 -08:00
Tony Wasserka 952e949e10 JIT: Clean up block tail writing code 2025-12-01 20:13:51 +01:00
Tony Wasserka 3d093d66fb JIT: Change relocation offset base to CodeBuffer start 2025-12-01 20:13:51 +01:00
Tony Wasserka 05fe2893c7 JIT: Emit relocation from IROP_THUNK 2025-12-01 20:13:51 +01:00
Tony Wasserka 6fc17294b6 JIT: Add relocation for guest RIP stored in jump thunks 2025-12-01 20:13:51 +01:00
Tony Wasserka 6607921bee JIT: Add relocation for guest RIP stored in block tail 2025-12-01 20:13:51 +01:00
Tony Wasserka 0b0793438f JIT: Add relocation for constants relative to the guest entrypoint 2025-12-01 20:13:51 +01:00
LC 3dd591e760 Merge pull request #5088 from Sonicadvance1/11
Linux: Disable io_uring
2025-12-01 14:05:56 -05:00
LC f290d2f899 Merge pull request #5090 from neobrain/refactor_3waycomp
IR: Replace hand-written operators with three-way comparison
2025-12-01 14:05:04 -05:00
Tony Wasserka a676ad7193 JIT: Add explicit padding to relocation descriptors
This ensures zero-initialization, which is required to make code cache
generation produce consistent results.

Also consolidated header fields.
2025-12-01 19:28:39 +01:00
Tony Wasserka 096c408ef6 IR: Replace hand-written operators with three-way comparison 2025-12-01 17:10:06 +01:00
Ryan Houdek 00b65f76b6 Linux: Disable io_uring
This allows passing around `epoll_event` structs which can't be
rewritten due to queues being managed by userspace.
2025-11-30 15:53:34 -08:00
Ryan Houdek 39fb266282 Merge pull request #5087 from neobrain/refactor_simpler_calls
OpcodeDispatcher: Simplify convoluted logic for computing call offsets
2025-11-28 10:10:22 -08:00
Ryan Houdek 3969d0ac78 Merge pull request #5086 from neobrain/fix_invalid_iterators
LookupCache: Fix use of invalidated iterators
2025-11-28 09:55:20 -08:00
Tony Wasserka 439c6bb3c0 OpcodeDispatcher: Simplify convoluted logic for computing call offsets 2025-11-28 11:32:11 +01:00
Tony Wasserka 5cedbf9d34 LookupCache: Fix use of invalidated iterators 2025-11-28 11:29:46 +01:00
Tony Wasserka 427b235eb5 Merge pull request #5071 from Sonicadvance1/2
Github: Add a steamrt4 builder
2025-11-28 08:59:46 +00:00
LC 92d5ba580f Merge pull request #5085 from Sonicadvance1/10
HostFeatures: Extend LRCPC2 errata to more CPUs
2025-11-28 00:41:50 -05:00
Ryan Houdek bc2f331c8b HostFeatures: Extend LRCPC2 errata to C1 Ultra/Premium 2025-11-27 21:29:17 -08:00
Ryan Houdek ca58aef676 HostFeatures: Extend LRCPC2 errata to V3AE 2025-11-27 21:18:51 -08:00
LC 379dc405f6 Merge pull request #5080 from Sonicadvance1/9
FEXInterpreter: Fixes crash with code maps
2025-11-27 20:42:20 -05:00
LC d74b5c42da Merge pull request #5079 from Sonicadvance1/8
FEXServer: Add support for a `wait_fd`
2025-11-27 20:41:22 -05:00
LC 6004971439 Merge pull request #5077 from Sonicadvance1/6
Scripts/InstallFEX: Fixes two issues
2025-11-27 20:39:07 -05:00
LC a12b8927bc Merge pull request #5076 from Sonicadvance1/5
Utils/WritePriorityMutex: Support being forkable
2025-11-27 20:38:37 -05:00
Ryan Houdek 57e23b289a Merge pull request #5081 from esullivan-nvidia/main
HostFeatures: Disable SupportsTSOImm9 for some CPUs
2025-11-27 16:22:40 -08:00
Tony Wasserka 8214ffccf0 Merge pull request #5083 from bylaws/wheiwofjsd
JIT: Fix indirect delinker branch distance
2025-11-27 15:02:47 +00:00
Billy Laws 547135dc2d JIT: Fix indirect delinker branch distance
This is in insts not bytes.
2025-11-27 14:42:56 +00:00
Ryan Houdek 9c72113161 Merge pull request #5070 from wcampbell-nv/cmdline
Reflect application changes to argv[0] in /proc/self/cmdline
2025-11-26 19:50:22 -08:00
esullivan 8cc967fa22 HostFeatures: Disable SupportsTSOImm9 for some CPUs
This change avoids using the LDAPUR instruction with CPUs that are know to be
impacted by an ARM CPU errata that results in poor performance.
2025-11-26 21:42:29 -06:00
Ryan Houdek 4bd30bb72d HostFeatures: Fixes bug in HostFeatures where simulator doesn't support new things
We now have a machine in CI that requires this.
2025-11-26 15:43:37 -08:00
Ryan Houdek 5eeb4dabbd Github: Add a steamrt4 builder
This ensures we don't break downstream projects.
2025-11-26 14:18:08 -08:00
Tony Wasserka a27c4b3860 Merge pull request #5078 from Sonicadvance1/7
CPUID: Fixes regression from #5033
2025-11-26 15:19:51 +00:00
Ryan Houdek dfee08f74f FEXInterpreter: Fixes crash with code maps
When the realpath of a program path can't be resolved, we weren't setting
the config option. This was cascading to be a crashing in codemaps where
it was unconditionally using the optional value (with assert checks),
and causing things to crash.

Pass in the path that can't be resolved to work around a crash in PV
that can happen.
2025-11-25 16:48:47 -08:00
Will Campbell 0e2629bdd4 Address review feedback 2025-11-25 14:09:29 -08:00
Ryan Houdek 05c8630b07 FEXServer: Add support for a wait_fd
This was a requested feature. To make sure that FEXServer is running and
managed by a parent process, we need to have a way to tell FEXServer to
keep alive without any FEX clients. The best way to do this is to pass
FEXServer a Pipe (like FEX does when a client starts it), but instead of
FEXServer signaling to FEXInterpreter that it's ready. FEXServer listens
to the pipe to see if the management process is still alive.

The expectation here is that the management process passes FEXServer the
read end of a pipe, and when the management software is done (or gets
killed by the kernel!) then the write end of the pipe is closed, and
FEXServer naturally closes (As long as there's no FEX processes
remaining).
2025-11-25 12:52:16 -08:00
Ryan Houdek 1cffe618d2 CPUID: Fixes #5033
This leaf changed to being non-constant on that PR since CPUID function
1h returns APICID now.
2025-11-25 11:25:36 -08:00
Ryan Houdek 98c7bb23b5 FEX/InstallFEX: Fixes issue of installing without software-properties-common
Checks to see if the package is installed first before trying to use it.
Fixes an issue where fresh users don't have this package installed and
the script fails.
2025-11-24 14:41:48 -08:00
Ryan Houdek 922853cee1 Scripts/InstallFEX: Fixes #4972
Makes sure that stderr output doesn't cause weird interactions with
FEXRootFSFetcher.
2025-11-24 14:40:56 -08:00
Tony Wasserka a251e61859 Merge pull request #5075 from Sonicadvance1/4
Config: Document the new `FEX_APP_CACHE_LOCATION` option
2025-11-24 20:24:59 +00:00
Ryan Houdek 9c19799023 Config: Document the new FEX_APP_CACHE_LOCATION option
I forgot to document this in the man page.
2025-11-24 12:07:54 -08:00
Ryan Houdek f423b110a8 Utils/WritePriorityMutex: Support being forkable
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.

This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
2025-11-24 11:50:20 -08:00
Ryan Houdek 6fd471e652 Merge pull request #5073 from discapes/unsquashfs-deco-fix
Support detecting unsquashfs>4.7.0 decompressors
2025-11-24 09:05:08 -08:00
Tony Wasserka 3d69029d33 Merge pull request #5072 from Sonicadvance1/3
Minor fixes
2025-11-24 11:24:07 +00:00
Miika Tuominen 2258f2e424 Support detecting unsquashfs>4.7.0 decompressors 2025-11-22 14:38:59 +02:00
Ryan Houdek cca5a68e20 SHMStats: Add missing header 2025-11-21 18:01:37 -08:00
Ryan Houdek da5c9bff68 Async: Add missing header. 2025-11-21 18:01:33 -08:00
Ryan Houdek 5b87f0699b pidof: Switch to using ranges 2025-11-21 18:01:28 -08:00
Will Campbell a67fe561a1 Reflect application changes to argv[0] in /proc/self/cmdline 2025-11-21 16:00:57 -08:00
Tony Wasserka e2f4065376 Merge pull request #5065 from Sonicadvance1/1
FEXCore/CodeCache: Move spin-loop over to a WFE loop
2025-11-21 14:21:45 +01:00
Ryan Houdek d3bf87f4f4 Merge pull request #4985 from Sonicadvance1/fex-atomic
Support (downstream) kernel-side unaligned atomic handling (The rebase sequel)
2025-11-20 17:09:29 -08:00
Ryan Houdek 6d351ec47f FEXCore/CodeCache: Moves spin-loop in to a WFE loop
Saves power and responds faster. Pass in the atomic to `WaitPred` with
the predicate checking if the buffer has been flushed yet. Same
behaviour as previous code but more efficient on our hardware.
2025-11-20 14:23:01 -08:00
Ryan Houdek 4dc1dd2511 FEXCore/Utils/SpinWaitLock: Adds WaitPred for waiting on a predicate
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.

We will need a predicated version for the next commit.
2025-11-20 14:23:01 -08:00
Ryan Houdek 5205ae40fa Merge pull request #5062 from pmatos/fix/address-size-handle
Implement address size modifier handling in CMPSOp and SCASOp
2025-11-20 12:19:22 -08:00
Ryan Houdek 32f1dcde7e Merge pull request #4906 from neobrain/feature_code_maps
CodeCache: Introduce code maps
2025-11-20 12:11:16 -08:00
Ryan Houdek 6772581c53 Arm64ec: Print a log when kernel unaligned atomics are used
Not having this in FEXInterpreter as it is too spammy in the general
case.
2025-11-20 12:05:55 -08:00
Ryan Houdek 387201815b Config: Add option to enable or disable kernel backpatchin on unaligned atomic
Default to enabled because this is the config we expect by default.
In the future will get some benchmarking in various games like
Assassin's Creed, and Call of Duty.
2025-11-20 12:04:59 -08:00
Billy Laws 24d61e1125 Windows: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Billy Laws 04259f031d FEXLoader: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Tony Wasserka 6403da3715 LinuxSyscalls: Shield code map FD from guest access
This prevents chromium/CEF from closing the FD.
2025-11-20 19:13:18 +01:00
Tony Wasserka 90cb76312c LinuxSyscalls: Implement code map writing for future code caching 2025-11-20 19:13:18 +01:00
Tony Wasserka b6cff01abb FEXServer: Add support for querying code maps
Managing code maps in FEXServer rather than in FEXInterpreter makes it
easier to handle multiple concurrent processes sharing code caches for
the main executable and libraries.
2025-11-20 19:13:18 +01:00
Tony Wasserka b34b711161 CodeCache: Add interfaces to describe and generate code maps
Code maps describe per-binary metadata used to generate caches. Currently,
this includes compiled block offsets and loaded shared libraries.
2025-11-20 19:13:18 +01:00
Ryan Houdek 709d767d61 FEX/Config: Allow override of cache location
This will be used.
2025-11-20 18:56:22 +01:00
Tony Wasserka e075916154 Config: Add interface to query cache directory 2025-11-20 18:56:22 +01:00
Paulo Matos de10154f29 instcountci: Implement address size modifier handling in CMPSOp and SCASOp for 64bits 2025-11-20 13:42:19 +01:00
Paulo Matos ba71e79e54 asm_tests: Implement address size modifier handling in CMPSOp and SCASOp 2025-11-20 13:42:19 +01:00
Paulo Matos 2cc70b8051 Implement address size modifier handling in CMPSOp and SCASOp for 64bits
A few games were generating "Can't handle adddress size".
I implemented 0x67 prefix handling for CMPSOp and SCASOP and improved
the error messages for the remainder. This will implement the address
modifier on 64bit systems, and keep issuing an error on 32bits.
2025-11-20 13:42:19 +01:00
Paulo Matos 9d965f94de asm_tests: Add 32-bit CMPS/SCAS tests without address size override 2025-11-20 13:42:19 +01:00
Ryan Houdek 40c2db4744 Merge pull request #5006 from bylaws/fasterrrrr
Introduce two-pass code invalidation model
2025-11-19 17:41:01 -08:00
Billy Laws 9a7285dca4 Windows: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws 8c00ac78b1 Linux: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws cf4478eeee LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-11-20 00:38:03 +00:00
Billy Laws 85c8e7f1bb fextl: Wrap tsl::robin_set 2025-11-20 00:38:03 +00:00
Billy Laws efd95efb40 FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-11-20 00:38:03 +00:00
Billy Laws 99ad7ea45c LookupCache: Drop unused state frame argument for delinker cbs 2025-11-20 00:38:03 +00:00
Ryan Houdek aba0c57f73 Merge pull request #5067 from pmatos/fix/Nasm3
Fix movzx instruction syntax
2025-11-19 14:01:15 -08:00
Paulo Matos 8c4f6b648e Fix movzx instruction syntax
nasm 2.16 was happy with it but it generates a bunch of errors in nasm3.
The generated binaries remain the same.
2025-11-19 15:04:16 +01:00
Ryan Houdek 3b83bdd88d Merge pull request #5066 from neobrain/fix_async_asserts
Async: Adapt precondition checks when receiving FDs
2025-11-19 01:45:41 -08:00
Tony Wasserka 8d71e08b44 Async: Strengthen precondition check when receiving FDs
The sender might provide all requested message bytes but no FD. The receiver
interface has no simple way of indicating this scenario yet, so just assert
out for now to ensure it never happens in the first place.

If needed, this can be changed to return a new error code to indicate partial
read in the future.
2025-11-19 09:48:23 +01:00
Tony Wasserka c31063a8ef Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-19 09:35:52 +01:00
Tony Wasserka 28d101f1dd Revert "Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark"
This cherry-picked an unfinished patch that wasn't intended for merging.
2025-11-19 09:35:08 +01:00
Ryan Houdek aaef344ae3 Merge pull request #5039 from Sonicadvance1/warkwarkwarkwarkwarkwark
FEX: Moves FEX thunk callback function generation to the frontend
2025-11-18 14:40:53 -08:00
Ryan Houdek 42d0324304 FEX: Moves FEX thunk callback function generation to the frontend
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.

Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.

Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.
2025-11-18 14:15:04 -08:00
Ryan Houdek 2a0019347a Thunks: Adds FEX Thunk callback to VDSO
This isn't necessary a VDSO, but it is /always/ mapped in to every
process. Use it as such.
2025-11-18 14:08:43 -08:00
Ryan Houdek e0305ea1b9 Merge pull request #5056 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEX/VDSO: Fixes symbol lookup
2025-11-18 12:51:13 -08:00
Ryan Houdek 105ff47ae3 FEX/VDSO: Fixes symbol lookup
Fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to .dynsym where gcc sticks them in to .symtab. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

Peeled out of #5039
2025-11-18 12:19:51 -08:00
Ryan Houdek 5ee190a41e Merge pull request #5060 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Config: Expose GetConv members
2025-11-17 10:45:02 -08:00
Ryan Houdek cf37617c25 Merge pull request #5011 from pmatos/feat/opt-memcpyf80
Refactoring of storing code in x87 opt. stack pass
2025-11-17 10:25:39 -08:00
Tony Wasserka 2e9c8f0f51 FEXCore/Config: Expose GetConv members
From working branch commit 8244ca1666796267ce25741cdf1103eef4f7539d
`Make cache generation aware of FEX configuration`
2025-11-17 10:21:51 -08:00
Ryan Houdek 06c2319851 Merge pull request #5061 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Common: Adds the ability to override HostFeatures registers by config.
2025-11-17 10:18:26 -08:00
Ryan Houdek da0668c7cc Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
Async: Move file descriptor checks from read_some() to read()
2025-11-17 10:17:44 -08:00
Ryan Houdek b34df334cb Merge pull request #5053 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection
2025-11-17 10:16:59 -08:00
Tony Wasserka 9e9f2ccae1 Merge pull request #5050 from neobrain/refactor_new_config_getter
Config: Refactor value getter interface
2025-11-17 18:58:50 +01:00
Tony Wasserka 15b8f75730 Config: Drop unneeded namespaces from StringArrayType 2025-11-17 18:46:35 +01:00
Tony Wasserka 9d6b9aa574 Config: Allow reading config values in arbitrary C++ expressions
FEX_CONFIG_OPT can only be used as a standalone statement, which is
inconvenient for config values that are only used once. The new functions
(e.g. Get_DUMPIR()) can be used in conditions or other expressions.
2025-11-17 18:46:02 +01:00
Tony Wasserka 64724886af Config: Replace macro-based config readers with a C++ template 2025-11-17 18:46:02 +01:00
Paulo Matos c088369f4a instcountci: Refactoring of storing code in x87 opt. stack pass 2025-11-17 10:14:29 +01:00
Paulo Matos 39dbf46422 Refactoring of storing code in x87 opt. stack pass
Enables memcpy optimization of 80bit floats on reduced precision.

Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
2025-11-17 10:14:29 +01:00
Paulo Matos 3c1b0bb917 Add instcountci tests for 80bit memcpy for x87 instructions 2025-11-17 10:14:29 +01:00
Ryan Houdek 2227170dbb FEXCore/Common: Adds the ability to override HostFeatures registers by config.
This is going to be necessary for the offline compiler work.
Also allow FEXGetConfig to print the same registers in the correct
format for easy fetching.
2025-11-16 16:43:01 -08:00
Ryan Houdek e08f421e1c Convert Async assert to logman assert 2025-11-16 14:56:30 -08:00
Tony Wasserka 42c58c5420 Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-16 14:54:46 -08:00
LC 4afbdd9afb Merge pull request #5055 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fixes alloca use-after-free in `RecvMMsg`
2025-11-14 20:26:05 -05:00
Ryan Houdek 06b9e13904 LinuxSyscalls: Fixes alloca use-after-free in RecvMMsg
Easy enough fix, thanks to @OFFTKP for pointing this out.
2025-11-14 16:27:27 -08:00
Tony Wasserka 29473b43cb LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection 2025-11-14 15:23:42 -08:00
Ryan Houdek 0427d48b98 Merge pull request #5033 from Sonicadvance1/warkwark
CPUID: Fixes APICID for processor count calculation.
2025-11-14 11:31:26 -08:00
Tony Wasserka b228746f1d FEXInterpreter: Fix incorrect value assignment
Previously this was overriding the cached value in the local Value object.
The global SilentLog variable never got updated, so a stale value would be
used.
2025-11-14 11:32:11 +01:00
Tony Wasserka 0be8485116 Merge pull request #5047 from antonkesy/fix_formatting
Align code with clang-format
2025-11-14 09:25:38 +01:00
Tony Wasserka 5d0279ff08 LibraryForwarding/gen: Tiny cleanup 2025-11-14 09:14:45 +01:00
Ryan Houdek 5d908d902c Merge pull request #5045 from lioncash/long
JIT: Handle long ADR/ADRP
2025-11-13 11:10:10 -08:00
Ryan Houdek 3a014f80f2 Merge pull request #5044 from antonkesy/clean_up_scripts
Scripts: Clean-up
2025-11-13 11:09:56 -08:00
Ryan Houdek fbefd7855c Merge pull request #5048 from neobrain/fix_fexconfig_string_lists
FEXConfig: Fix string list handling
2025-11-13 11:07:50 -08:00
Tony Wasserka 11f9135be6 FEXConfig: Fix string list handling 2025-11-13 17:33:21 +01:00
Lioncache 7bb0ce810e JIT: Expand LongAddressGen() to handle movz+movk sequence
This is only ever used on the path where we'd want to handle something
like this (in EmitEntryPoint()), so we can just extend the long handler
type instead of introducing a new type to handle this.
2025-11-13 11:16:34 -05:00
Lioncache 993b832771 JIT: Handle long ADR/ADRP
Wires up the long address handler into the ADR/ADRP restart
handlers.
2025-11-13 09:40:39 -05:00
Anton Kesy 013ac1e627 Align code with clang-format
Automatically done by running:
`find . \( -path './External' -prune \) -o \
  \( -iname '*.cc' -o -iname '*.cpp' -o -iname '*.hpp' -o \
     -iname '*.h' -o -iname '*.c' \) -print | \
  xargs clang-format --style=file -i`
2025-11-13 13:29:49 +01:00
Anton Kesy d7977a02fa remove semicolon 2025-11-12 21:45:09 +01:00
LC 73a32ff22c Merge pull request #5043 from antonkesy/fix_typos
Docs: fix typo
2025-11-12 15:28:24 -05:00
Anton Kesy ee4ae5390b move function comment inside function 2025-11-12 21:22:26 +01:00
Anton Kesy f71db11035 remove excess whitespaces 2025-11-12 21:22:08 +01:00
Anton Kesy 9497288b97 remove unused imports 2025-11-12 21:21:55 +01:00
Anton Kesy a00260d801 remove unused variable 2025-11-12 21:21:28 +01:00
Anton Kesy cad48e07e4 fix comment indentation 2025-11-12 21:21:05 +01:00
Anton Kesy d91e8a4278 docs: fix typo 2025-11-12 21:07:18 +01:00
LC faf74eee90 Merge pull request #5037 from Sonicadvance1/warkwarkwarkwark
FEXCore: Fixes JITGuardPage calculation in a threaded environment
2025-11-12 15:00:00 -05:00
Ryan Houdek 0b52e1cd14 Merge pull request #5042 from neobrain/fix_base_inference
LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
2025-11-12 11:18:58 -08:00
Tony Wasserka b62890f136 LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
At runtime, glxtest is mapped as follows:
0x000055fd9a030000 0x000055fd9a034000 0x4000  0x0     r--p  glxtest
0x000055fd9a034000 0x000055fd9a038000 0x4000  0x3000  r-xp  glxtest
0x000055fd9a038000 0x000055fd9a039000 0x1000  0x6000  rw-p  glxtest
0x000055fd9a039000 0x000055fd9a03a000 0x1000  0x6000  rw-p  glxtest

The problem here is that the last two sections can't be distinguished solely
by their mmap parameters. This would cause the wrong base address to be
inferred for the last mapping. To fix this, we can be more permissive by
allowing multiple candidates to be returned.

In practice, this only affects non-code sections, so it's not a big issue
either way.
2025-11-12 19:54:45 +01:00
Tony Wasserka de1d37eef8 Merge pull request #5041 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwark
CodeEmitter: Removes a few spurious asserts
2025-11-12 14:34:16 +01:00
Ryan Houdek a57c557485 CodeEmitter: Removes a few spurious asserts
These are handled with restart.
2025-11-11 18:25:21 -08:00
LC 3c9f6c845b Merge pull request #5040 from Sonicadvance1/warkwarkwarkwarkwarkwarkwark
FEXCore: Remove usage of "remote atomic" xor
2025-11-11 21:20:18 -05:00
Ryan Houdek ff25e9a92e FEXCore: Remove usage of "remote atomic" xor
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
2025-11-11 17:16:03 -08:00
Ryan Houdek 6a60f72a9e FEXCore: Fixes JITGuardPage calculation in a threaded environment
While this worked great for the singular unit test. I remembered thatour
pool allocator returns the minimum working size asked for but will
return larger sizes if exact fitment couldn't occur.

Because we are dealing with guard pages, we need to return the full
buffer size to the "client" so they can tell the frontend where the
guard page actually lives. Otherwise the JIT will tell the frontend the
guard page is at the end of the requested size, blow past the limit,
and fault in a completely different location.

With a bit of logging I saw in a multithreaded environment that we were
basically always getting a larger requested buffer while Steam was
starting up.
2025-11-11 12:31:38 -08:00
Ryan Houdek 94b690df43 unittests/FEXLinuxTests: Adds cpu core count test to cpuid
Ensures cpuid core counts are reported correctly.
2025-11-11 11:20:26 -08:00
Ryan Houdek 94edbc3436 CPUID: Stop accidentally exposing the HTT bit
We don't support this.
2025-11-11 11:20:26 -08:00
Ryan Houdek 5eab1e559a CPUID: Fixes APICID for processor count calculation.
Primary fix here is returning the current CPU index in function 01h.
Intel Quartus uses this alongside affinity setting to check if all cores
can be used for its calculation. Since we had hardcoded apicid 0 here,
it assumed to only have one core and never generated worker threads.

Additional fix for apicid size. This is the size of the bitmask required
for apic ids, we weren't calculating this correctly at all. This mask is
a "maximum" number of APICs that the CPU reserves in power of two.
Say the core supports 256 APICs, but the processor only supports 16, or
any other combination.
2025-11-11 11:20:26 -08:00
Ryan Houdek 53db3ad6f2 Merge pull request #5036 from neobrain/fix_thunkgen_glibcxx_debug
CMake: Disable libstdc++'s debug mode when compiling thunkgen
2025-11-11 09:14:53 -08:00
Tony Wasserka 581f3263ed CMake: Disable libstdc++'s debug mode when compiling thunkgen
This allows the rest of the project to use _GLIBCXX_DEBUG.
2025-11-11 17:26:37 +01:00
LC 1e3c642be6 Merge pull request #5035 from pmatos/fix/gradual-mem-growth
Use gradual memory growth
2025-11-11 09:00:27 -05:00
LC 22c3cd553f Merge pull request #5034 from Sonicadvance1/warkwarkwark
FEXCore/Win32: Move WritePriorityMutex away from SRWLock
2025-11-11 08:59:29 -05:00
Paulo Matos e5743f8dae Use gradual memory growth
Use min instead of max, otherwise we are always using `MAX_STATS_SIZE`.
2025-11-11 09:23:40 +01:00
Ryan Houdek bddc2f227d FEXCore/Win32: Move WritePriorityMutex away from SRWLock
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.

Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.

While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.

Dark Souls Remastered before:
```
  $RDLck Time: 4.531100 ms/second (0.04 percent)
  $WRLck Time: 2.122560 ms/second (0.02 percent)
```

after:
```
  $RDLck Time: 1.441620 ms/second (0.01 percent)
  $WRLck Time: 0.963720 ms/second (0.01 percent)
```
2025-11-10 17:43:47 -08:00
Ryan Houdek 686c04ea93 Win32: IMplement Wake/Wait by address 2025-11-10 17:43:40 -08:00
Ryan Houdek b38369199e Merge pull request #4893 from Sonicadvance1/long_long_codebuffer_pages
FEX: Implements support for JIT CodeBuffer guard page restart
2025-11-10 13:48:18 -08:00
Ryan Houdek e862c904a9 FEX: Implements support for JIT CodeBuffer guard page restart
When the JIT CodeBuffer overflows, we will now catch accesses to the
guard page and longjump while restarting the JIT with a larger buffer
request.

Fixes #4877
2025-11-10 11:55:21 -08:00
Ryan Houdek 43d9384b1c FEXCore/JIT: Add a pool allocator that understands a guard page
The size asked for has its final page guarded. It's up to the code
asking for allocations to ensure it never uses the final page if
necessary.
2025-11-10 11:54:25 -08:00
Ryan Houdek cb9af0b86a SignalDelegator: Split out SIGSEGV handler
This needs to run before the TestCodeHarness's frontend handler.
2025-11-10 11:53:40 -08:00
Ryan Houdek b8c17a843c ArchHelpers: Adds helper to get pointers to PC and FPRs 2025-11-07 15:33:47 -08:00
Ryan Houdek 7ad7f181d7 FEXCore/LongJump: Add a way to manually load from a longjump
The frontends will need this when loading a longjump buffer in to a
context.
2025-11-07 15:33:47 -08:00
Ryan Houdek eb0bf55033 FEXCore/JIT: Move the JIT long jump buffer to internalthreadstate
This will be a TLS variable that needs to be read by the frontend.
2025-11-07 15:33:47 -08:00
Ryan Houdek f4e3e4ad30 Merge pull request #5031 from lioncash/catch
Externals: Update catch2 from 3.5.3 to 3.11.0
2025-11-07 10:40:49 -08:00
Lioncache 5ae82410cc Externals: Update catch2 from 3.5.3 to 3.11.0
Updates it to the most recent release.
2025-11-07 08:31:04 -05:00
Ryan Houdek 2febb524e9 Merge pull request #5028 from lioncash/fmtup
Externals: Update fmt to 12.1.0
2025-11-06 10:04:15 -08:00
Lioncache b9e452133c Externals: Update fmt to 12.1.0
Keeps fmt updated to its latest release.
2025-11-06 10:17:22 -05:00
LC 747ea0a1f7 Merge pull request #5027 from Sbte/pr/xxhash
Update xxhash to v0.8.3
2025-11-06 07:38:56 -05:00
Sven Baars f8c52ca34a Update xxhash to v0.8.3 2025-11-06 11:53:55 +01:00
186 changed files with 4907 additions and 1334 deletions

No files matched your search

+79
View File
@@ -0,0 +1,79 @@
name: steamrt4 build
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
jobs:
steamrt4_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, distrobox]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name : submodule checkout
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build
cmake -E make_directory ${{runner.workspace}}/build
# Setup everything required.
- name : distrobox setup
run: |
distrobox create -Y -i registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306 steamrt4 || true
distrobox upgrade steamrt4
distrobox enter --name steamrt4 -- sudo apt-get install -y \
git cmake ninja-build ccache \
lld clang \
libclang-dev llvm-dev \
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross \
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- name: Create Build Environment
run: distrobox enter --name steamrt4 -- cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
shell: bash
working-directory: ${{runner.workspace}}/build
run: distrobox enter --name steamrt4 -- cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE
- name: install
working-directory: ${{runner.workspace}}/build
shell: bash
env:
DESTDIR: ${{runner.workspace}}/install
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE -t install
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
overwrite: true
name: steamrt4_steampipe_depot
path: ${{runner.workspace}}/install/*
retention-days: 1
compression-level: 9
+14 -2
View File
@@ -33,6 +33,7 @@ option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling ca
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use for the FEXCore profiler (gpuvis, tracy)")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
option(BUILD_STEAM_SUPPORT "Builds FEX for integration into Steam" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
@@ -64,6 +65,10 @@ if (NOT MINGW_BUILD)
endif()
endif()
if (BUILD_STEAM_SUPPORT)
add_definitions(-DFEX_STEAM_SUPPORT=1)
endif()
if (ENABLE_FEXCORE_PROFILER)
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
@@ -476,13 +481,16 @@ add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
if (_M_ARM_64 AND NOT MINGW_BUILD)
if (_M_ARM_64 AND NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
# Binfmt_misc files must be installed prior to Source/ installs
add_subdirectory(Data/binfmts/)
endif()
add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
if (NOT BUILD_STEAM_SUPPORT)
add_subdirectory(Data/AppConfig/)
endif()
# Install the ThunksDB file
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS ${CMAKE_CURRENT_SOURCE_DIR}/Data/*.json)
@@ -582,6 +590,10 @@ if (BUILD_THUNKS)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
if (BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
set(FEX_VERSION_MAJOR "0")
set(FEX_VERSION_MINOR "0")
set(FEX_VERSION_PATCH "0")
+19 -11
View File
@@ -38,8 +38,6 @@ public:
[[nodiscard]] BranchEncodeSucceeded adr(ARMEmitter::Register rd, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
if (IsADRRange(Imm)) {
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
@@ -73,7 +71,6 @@ public:
[[nodiscard]] BranchEncodeSucceeded adrp(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) {
constexpr uint32_t Op = 0b1001'0000 << 24;
@@ -103,16 +100,22 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>());
const auto SLocation = reinterpret_cast<int64_t>(Label->Location);
const auto ULocation = std::bit_cast<uint64_t>(SLocation);
const int64_t Imm = SLocation - (GetCursorAddress<int64_t>());
const auto UImm = std::bit_cast<uint64_t>(Imm);
if (IsADRRange(Imm)) {
// If the range is in ADR range then we can just use ADR.
return adr(rd, Label);
} else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
}
if (IsADRPRange(Imm)) {
const int64_t ADRPImm = (SLocation & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
// If the range is in the ADRP range then we can use ADRP.
bool NeedsOffset = !IsADRPAligned(reinterpret_cast<uint64_t>(Label->Location));
uint64_t AlignedOffset = reinterpret_cast<uint64_t>(Label->Location) & 0xFFFULL;
const bool NeedsOffset = !IsADRPAligned(ULocation);
const uint64_t AlignedOffset = ULocation & 0xFFFULL;
// First emit ADRP
adrp(rd, ADRPImm >> 12);
@@ -125,14 +128,19 @@ public:
return BranchEncodeSucceeded::Success;
}
// Can't encode.
return BranchEncodeSucceeded::Failure;
// Stinky path, we need to load the address as a sequence of movz+movk+movk
movz(ARMEmitter::Size::i64Bit, rd, (UImm >> 32) & 0xFFFF, 32);
movk(ARMEmitter::Size::i64Bit, rd, (UImm >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, rd, UImm & 0xFFFF);
return BranchEncodeSucceeded::Success;
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::LONG_ADDRESS_GEN});
// Emit a register index and a nop. These will be backpatched.
// Emit a register index and two nops. These will be backpatched.
dc32(rd.Idx());
nop();
nop();
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
-1
View File
@@ -301,7 +301,6 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded tbnz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0111 << 24;
+22 -16
View File
@@ -741,38 +741,44 @@ public:
break;
}
case ForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
const auto* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
const auto ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
const auto ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
const auto ImmInstThree = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[2]);
const auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
const auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
if (IsADRRange(ImmInstThree)) {
// If within ADR range from the third instruction, then we can emit NOP+NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
} else if (IsADRPRange(ImmInstOne)) {
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstThree) & 0x7FFF);
} else if (IsADRPRange(ImmInstTwo)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
// We can emit nop + nop + adrp
nop();
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstThree >> 12) & 0x7FFF);
} else {
// Not aligned, need nop + adrp + add
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
} else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstTwo & 0xFFF);
}
} else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
// Stinky path, we need to emit a movz+movk+movk sequence.
movz(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 32) & 0x7FFF, 32);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne) & 0xFFFF);
}
SetCursorOffset(OriginalOffset);
+1 -1
+1 -1
+1 -1
+5 -3
View File
@@ -74,9 +74,11 @@ add_compile_options($<$<COMPILE_LANGUAGE:CXX>:-fno-strict-aliasing> $<$<COMPILE_
add_subdirectory(Source/)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
if (NOT BUILD_STEAM_SUPPORT)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
endif()
if (BUILD_TESTING)
add_subdirectory(unittests/)
+9
View File
@@ -200,6 +200,15 @@ def print_man_environment_tail():
],
"''", True)
print_man_env_option(
"APP_CACHE_LOCATION",
[
"Allows the user to override where FEX stores and loads cache files",
"By default FEX will look in $XDG_CACHE_HOME/fex-emu/ or $HOME/.cache/fex-emu/",
"This will override the full path, trailing forward-slash is expected to exist",
],
"''", True)
def print_man_header():
header ='''.Dd {0}
.Dt FEX
+5 -4
View File
@@ -31,7 +31,6 @@ set (SRCS
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
@@ -202,8 +201,10 @@ add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_MAN_NAME}"
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
if (NOT BUILD_STEAM_SUPPORT)
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
endif()
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
@@ -290,7 +291,7 @@ AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
if (NOT MINGW_BUILD)
if (NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
install(TARGETS ${PROJECT_NAME}_shared
LIBRARY
DESTINATION ${CMAKE_INSTALL_LIBDIR}
+19 -14
View File
@@ -30,14 +30,14 @@ class Context;
}
namespace FEXCore::Config {
namespace DefaultValues {
namespace detail {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) const P(type) P(enum) = P(default);
#define OPT_STR(group, enum, json, default) const std::string_view P(enum) = P(default);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#define OPT_STRENUM(group, enum, json, default) const uint64_t P(enum) = FEXCore::ToUnderlying(P(default));
#include <FEXCore/Config/ConfigValues.inl>
} // namespace DefaultValues
} // namespace detail
enum Paths {
PATH_DATA_DIR_LOCAL = 0,
@@ -134,7 +134,7 @@ public:
void Load();
template<typename T>
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<DefaultValues::Type::StringArrayType, T>)
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<StringArrayType, T>)
std::optional<T> GetConv(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
@@ -142,7 +142,7 @@ public:
}
const auto& Value = it->second;
LOGMAN_THROW_A_FMT(!std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(!std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
if (std::holds_alternative<T>(Value)) [[likely]] {
return std::get<T>(Value);
@@ -165,7 +165,7 @@ public:
private:
void MergeConfigMap(const LayerOptions& Options);
void MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value);
void MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value);
};
void MetaLayer::Load() {
@@ -181,7 +181,7 @@ void MetaLayer::Load() {
}
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value) {
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
@@ -193,7 +193,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](const DefaultValues::Type::StringArrayType& Value) {
const auto AddToMap = [&LookupMap](const StringArrayType& Value) {
for (const auto& EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == fextl::string::npos) {
@@ -209,7 +209,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
}
};
AddToMap(std::get<DefaultValues::Type::StringArrayType>(MetaEnvironment->second));
AddToMap(std::get<StringArrayType>(MetaEnvironment->second));
AddToMap(Value);
// Now with the two layers merged in the map
@@ -225,8 +225,8 @@ void MetaLayer::MergeConfigMap(const LayerOptions& Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto& it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV || it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<DefaultValues::Type::StringArrayType>(it.second));
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<StringArrayType>(it.second));
} else {
OptionMap.insert_or_assign(it.first, it.second);
}
@@ -423,7 +423,7 @@ bool Exists(ConfigOption Option) {
return Meta->OptionExists(Option);
}
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
std::optional<StringArrayType*> All(ConfigOption Option) {
return Meta->All(Option);
}
@@ -436,6 +436,12 @@ std::optional<T> GetConv(ConfigOption Option) {
return Meta->GetConv<T>(Option);
}
template std::optional<bool> GetConv(ConfigOption Option);
template std::optional<uint8_t> GetConv(ConfigOption Option);
template std::optional<int32_t> GetConv(ConfigOption Option);
template std::optional<uint32_t> GetConv(ConfigOption Option);
template std::optional<uint64_t> GetConv(ConfigOption Option);
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
@@ -491,13 +497,12 @@ template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t De
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List) {
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<DefaultValues::Type::StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option,
DefaultValues::Type::StringArrayType* List);
template void Value<StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
} // namespace FEXCore::Config
@@ -16,6 +16,13 @@
"Maximum number of instruction to store in a block"
]
},
"EnableCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable the code caching subsystem"
]
},
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
@@ -94,6 +101,13 @@
"Desc": [
"Scales the cycle counter on systems that have low frequencies."
]
},
"CPUFeatureRegisters": {
"Type": "str",
"Default": "",
"Desc": [
"Allows overriding cpu feature flags for manual testing"
]
}
},
"Emulation": {
@@ -429,6 +443,13 @@
"This is required to ensure a split-lock doesn't tear inside the process"
]
},
"KernelUnalignedAtomicBackpatching": {
"Type": "bool",
"Default": "true",
"Desc": [
"When the kernel unaligned atomic handler is enabled, use backpatching to reduce kernel context switches."
]
},
"VolatileMetadata": {
"Type": "bool",
"Default": "true",
+41 -9
View File
@@ -4,7 +4,6 @@
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -62,7 +61,7 @@ struct CustomIRResult {
, Data(Data) {}
};
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class CodeCache : public AbstractCodeCache {
@@ -73,12 +72,34 @@ public:
ContextImpl& CTX;
bool IsGeneratingCache = false;
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
void LoadData(Core::InternalThreadState&, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
void InitiateCacheGeneration() override {
IsGeneratingCache = true;
}
/**
* Applies a set of FEX relocations to the given code section.
*
* FEX relocations describe runtime-dependencies of FEX-generated code.
* When loading a code cache, they are used to move cached code to the
* dynamically chosen base address of the guest binary.
*
* Conversely, relocations are applied in reverse when writing code caches
* to ensure consistency across generation runs.
*
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
@@ -155,10 +176,20 @@ public:
return CodeCache;
}
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter> Writer) override {
CodeMapWriter = std::move(Writer);
}
void FlushAndCloseCodeMap() override {
if (CodeMapWriter) {
CodeMapWriter.reset();
}
}
void OnCodeBufferAllocated(const std::shared_ptr<CPU::CodeBuffer>&) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start,
uint64_t Length) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -225,14 +256,12 @@ public:
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CodeCache CodeCache;
fextl::unique_ptr<CodeMapWriter> CodeMapWriter;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl(const FEXCore::HostFeatures& Features);
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, const FEXCore::LookupCacheWriteLockToken& lk);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
// This is used as a replacement for the SMC writes in the mono callsite backpatcher that avoids atomic operations
@@ -269,7 +298,7 @@ public:
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator {"FEXMem_OpDispatcher"};
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator {"FEXMem_Frontend"};
FEXCore::Utils::PooledAllocatorVirtual CPUBackendAllocator {"FEXMem_CPUBackend"};
FEXCore::Utils::PooledAllocatorVirtualWithGuard CPUBackendAllocator {"FEXMem_CPUBackend"};
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const {
@@ -348,5 +377,8 @@ private:
bool MonoDetected = false;
std::atomic<uint64_t> MonoBackpatcherBlock;
std::mutex CodeBufferListLock;
fextl::vector<std::weak_ptr<CPU::CodeBuffer>> CodeBufferList;
};
} // namespace FEXCore::Context
@@ -105,9 +105,12 @@ constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public ARMEmitter::Emitter {
protected:
public:
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
protected:
FEXCore::Context::ContextImpl* EmitterCTX;
std::span<const ARMEmitter::Register> StaticRegisters {};
@@ -117,8 +120,6 @@ protected:
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
// Correlate an ARM register back to an x86 register index.
+1 -1
View File
@@ -400,7 +400,7 @@ namespace CPU {
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(*Buffer);
OnCodeBufferAllocated(Buffer);
return Buffer;
}
+2 -2
View File
@@ -81,7 +81,7 @@ namespace CPU {
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(CodeBuffer&) {};
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
@@ -161,7 +161,7 @@ namespace CPU {
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) = 0;
virtual void ClearCache() {}
+9 -8
View File
@@ -14,6 +14,7 @@ $end_info$
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Syscalls.h>
@@ -443,10 +444,10 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
Res.eax = FAMILY_IDENTIFIER;
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(0 << 24); // Local APIC ID
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(GetCPUID() << 24); // Local APIC ID
Res.ecx = (1 << 0) | // SSE3
(CTX->HostFeatures.SupportsPMULL_128Bit << 1) | // PCLMULQDQ
@@ -509,7 +510,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
(1 << 25) | // SSE
(1 << 26) | // SSE2
(0 << 27) | // Self Snoop
(1 << 28) | // Max APIC IDs reserved field is valid
(0 << 28) | // (HTT) Max APIC IDs reserved field is valid
(1 << 29) | // Thermal monitor
(0 << 30) | // Reserved
(0 << 31); // Pending break enable
@@ -1096,9 +1097,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) con
(CTX->HostFeatures.SupportsCLZERO << 0); // CLZERO support
uint32_t CoreCount = Cores - 1;
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
((uint32_t)std::log2(CoreCount + 1) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
(std::bit_ceil(Cores) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
return Res;
}
+1 -1
View File
@@ -277,7 +277,7 @@ private:
// 0: Highest function parameter and ID
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 1: Processor info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 2: Cache and TLB info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 3: Serial Number(previously), now reserved
+347 -2
View File
@@ -1,12 +1,213 @@
// SPDX-License-Identifier: MIT
#include <Interface/Context/Context.h>
#include "Utils/SpinWaitLock.h"
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
#include <Interface/Core/JIT/Relocations.h>
#include <Interface/Core/LookupCache.h>
#include <FEXCore/Core/Thunks.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <git_version.h>
#include <xxhash.h>
#include <fstream>
namespace FEXCore {
#if __clang_major__ < 16
ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map, uint64_t FileId, fextl::string Filename)
: SourcecodeMap(std::move(Map))
, FileId(FileId)
, Filename(Filename) {}
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
std::string_view base_filename = FHU::Filesystem::GetFilename(std::string_view {MainExecutable.Filename});
if (FileId != 0xffff'ffff'ffff'ffff) {
return fextl::fmt::format("{}-{:016x}{}", base_filename, MainExecutable.FileId, AddNombSuffix ? "-nomb" : "");
}
return "";
}
fextl::map<CodeMapFileId, CodeMap::ParsedContents> CodeMap::ParseCodeMap(std::ifstream& File) {
fextl::map<CodeMapFileId, CodeMap::ParsedContents> Ret;
while (true) {
Entry Entry;
File.read(reinterpret_cast<char*>(&Entry), sizeof(Entry));
if (!File) {
break;
}
if (Entry.FileId == LoadExternalLibrary.FileId && Entry.BlockOffset == LoadExternalLibrary.BlockOffset) {
ExternalLibraryInfo Info;
File.read(reinterpret_cast<char*>(&Info), sizeof(Info));
fextl::string Filename;
std::getline(File, Filename, '\0');
// Align to 4-byte boundary
char Null[4];
File.read(Null, AlignUp(Filename.size() + 1, 4) - Filename.size() - 1);
if (!File) {
break;
}
Ret[Info.ExternalFileId].Filename = std::move(Filename);
} else if (Entry.FileId == SetExecutableFileId {}.Marker.FileId && Entry.BlockOffset == SetExecutableFileId {}.Marker.BlockOffset) {
CodeMapFileId ExecutableFileId;
File.read(reinterpret_cast<char*>(&ExecutableFileId), sizeof(ExecutableFileId));
if (!File) {
break;
}
Ret[ExecutableFileId].IsExecutable = true;
} else {
if (!Ret.contains(Entry.FileId)) {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
} else {
Ret[Entry.FileId].Blocks.insert(Entry.BlockOffset);
}
}
if (!File) {
break;
}
}
return Ret;
}
CodeMapWriter::CodeMapWriter(CodeMapOpener& Opener, bool OpenEagerly)
: Buffer(4096)
, FileOpener(Opener) {
if (OpenEagerly) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
}
CodeMapWriter::~CodeMapWriter() {
if (CodeMapFD.value_or(-1) != -1) {
Flush(BufferOffset);
close(*CodeMapFD);
}
}
bool CodeMapWriter::IsWriteEnabled(const ExecutableFileSectionInfo& Section) {
if (CodeMapFD == -1) {
return false;
}
// PV libraries can't yet be read by FEXServer, so skip dumping them
if (Section.FileInfo.Filename.starts_with("/run/pressure-vessel")) {
return false;
}
if (CodeMapFD) {
return true;
}
// Acquire mutex and re-check CodeMapFD to avoid race conditions
auto lk = std::unique_lock {Mutex};
if (!CodeMapFD) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
return CodeMapFD != -1;
}
void CodeMapWriter::Flush(size_t Offset) {
// Acquire exclusive lock and flush circular buffer
std::unique_lock Lock {Mutex};
Flush(Offset, Lock);
}
void CodeMapWriter::Flush(size_t Offset, std::unique_lock<std::shared_mutex>&) {
write(*CodeMapFD, Buffer.data(), Offset);
BufferOffset = 0;
}
void CodeMapWriter::AppendBlock(const FEXCore::ExecutableFileSectionInfo& SectionInfo, uint64_t BlockEntry) {
if (!IsWriteEnabled(SectionInfo)) {
return;
}
BlockEntry -= SectionInfo.FileStartVA;
if (BlockEntry > std::numeric_limits<uint32_t>::max()) {
ERROR_AND_DIE_FMT("Cannot write code map");
}
// Register new library if not already known
bool NewLibraryLoad = false;
{
// Check prior registration with shared lock
std::shared_lock Lock {Mutex};
NewLibraryLoad = !KnownFileIds.contains(SectionInfo.FileInfo.FileId);
}
if (NewLibraryLoad) {
// Register to map with exclusive lock
std::unique_lock Lock {Mutex};
NewLibraryLoad &= KnownFileIds.insert(SectionInfo.FileInfo.FileId).second;
}
if (NewLibraryLoad) {
// Add entry to code map
AppendLibraryLoad(SectionInfo.FileInfo);
}
// Register the actual code block
CodeMap::Entry DataEntry {SectionInfo.FileInfo.FileId, static_cast<uint32_t>(BlockEntry)};
AppendData(std::as_bytes(std::span {&DataEntry, 1}));
}
void CodeMapWriter::AppendLibraryLoad(const FEXCore::ExecutableFileInfo& FileInfo) {
// See CodeMap::ExternalLibraryInfo
auto ExternalFileId = FileInfo.FileId;
auto TotalSize = AlignUp(sizeof(CodeMap::LoadExternalLibrary) + sizeof(ExternalFileId) + FileInfo.Filename.size() + 1, 4);
const auto Data = reinterpret_cast<char*>(alloca(TotalSize));
auto WritePtr = std::copy_n(reinterpret_cast<const char*>(&CodeMap::LoadExternalLibrary), sizeof(CodeMap::LoadExternalLibrary), Data);
WritePtr = std::copy_n(reinterpret_cast<const char*>(&ExternalFileId), sizeof(ExternalFileId), WritePtr);
WritePtr = std::copy(FileInfo.Filename.begin(), FileInfo.Filename.end(), WritePtr);
std::fill(WritePtr, Data + TotalSize, 0);
AppendData(std::as_bytes(std::span {Data, TotalSize}));
}
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo) {
CodeMap::SetExecutableFileId Data {.ExecutableFileId = FileInfo.FileId};
AppendData(std::span {reinterpret_cast<const std::byte*>(&Data), sizeof(Data)});
}
void CodeMapWriter::AppendData(std::span<const std::byte> Data) {
std::shared_lock Lock {Mutex};
auto Offset = BufferOffset.fetch_add(Data.size_bytes());
if (Offset + Data.size_bytes() > Buffer.size()) {
// Acquire exclusive lock and flush the buffer.
// Under heavy pressure, multiple threads may observe an exhausted buffer simultaneously.
// The thread with the last in-bounds Offset is responsible for flushing the buffer.
Lock.unlock();
bool IsResponsibleForFlush = false;
{
std::unique_lock ExclusiveLock {Mutex};
IsResponsibleForFlush = (Offset <= Buffer.size());
if (IsResponsibleForFlush) {
Flush(Offset, ExclusiveLock);
}
}
if (!IsResponsibleForFlush) {
// Wait for the buffer to be flushed on the responsible thread
Utils::SpinWaitLock::WaitPred<std::less_equal<>, size_t>(reinterpret_cast<size_t*>(&BufferOffset), Buffer.size());
}
AppendData(Data);
return;
}
memcpy(&Buffer.at(Offset), Data.data(), Data.size_bytes());
}
} // namespace FEXCore
namespace FEXCore::Context {
@@ -15,12 +216,156 @@ CodeCache::CodeCache(ContextImpl& CTX_)
: CTX(CTX_) {}
CodeCache::~CodeCache() = default;
uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
if (Filename.empty()) {
return 0xffff'ffff'ffff'ffff;
}
// For now, we just use the file path as an identifier.
// TODO: Ensure the hash is unique enough to distinguish executables while remaining independent of the installation location
return XXH3_64bits(Filename.data(), Filename.size());
}
struct CodeCacheHeader {
char Magic[4] = {'F', 'X', 'C', 'C'};
uint32_t FormatVersion = 1;
char FEXVersion[8] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
uint32_t CodeBufferSize;
uint32_t NumRelocations;
uint64_t SerializedBaseAddress;
// TODO: Consider including information from LookupCache.BlockLinks
};
void CodeCache::LoadData(Core::InternalThreadState& Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& GuestRIPLookup) {
// TODO
}
template<typename T>
static constexpr auto IsOrderedContainer(const T&) -> std::false_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::map<T...>&) -> std::true_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::set<T...>&) -> std::true_type;
bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const ExecutableFileSectionInfo& SourceBinary, uint64_t SerializedBaseAddress) {
// TODO
auto CodeBuffer = CTX.GetLatest();
auto& LookupCache = *Thread.LookupCache->Shared;
auto Relocations = Thread.CPUBackend->TakeRelocations(SourceBinary.FileStartVA);
// Write file header
CodeCacheHeader header;
memcpy(&header.FEXVersion[0], GIT_SHORT_HASH, strlen(GIT_SHORT_HASH));
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
// Dump guest<->host block mappings
{
// Cache contents must be deterministic, so copy the unordered block list and then sort by key
static_assert(!decltype(IsOrderedContainer(LookupCache.BlockList))::value, "Already deterministic; drop temporary container");
fextl::vector<std::pair<uint64_t, const GuestToHostMap::BlockEntry*>> BlockList;
BlockList.reserve(LookupCache.BlockList.size());
for (auto& [Guest, BlockEntry] : LookupCache.BlockList) {
static_assert(sizeof(Guest) == 8, "Breaking change in code cache data layout");
BlockList.emplace_back(Guest, &BlockEntry);
}
std::ranges::sort(BlockList);
for (auto [Guest, Host] : BlockList) {
static_assert(sizeof(Host->HostCode) == 8, "Breaking change in code cache data layout");
static_assert(sizeof(Host->CodePages[0]) == 8, "Breaking change in code cache data layout");
Guest -= SourceBinary.FileStartVA;
::write(fd, &Guest, sizeof(Guest));
uint64_t HostCode = Host->HostCode - reinterpret_cast<uintptr_t>(CodeBuffer->Ptr);
::write(fd, &HostCode, sizeof(HostCode));
uint64_t NumCodePages = Host->CodePages.size();
::write(fd, &NumCodePages, sizeof(NumCodePages));
LOGMAN_THROW_A_FMT(std::ranges::is_sorted(Host->CodePages), "Code pages aren't sorted");
for (auto CodePage : Host->CodePages) {
CodePage -= SourceBinary.FileStartVA;
::write(fd, &CodePage, sizeof(CodePage));
}
}
}
// Dump relocations
static_assert(sizeof(Relocations[0]) == 48, "Breaking change in code cache data layout");
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
}
// Dump the host code (relocated for position-independent serialization)
std::vector CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Dump code pages
static_assert(decltype(IsOrderedContainer(LookupCache.CodePages))::value, "Non-deterministic data source");
for (auto& [Page, Entrypoints] : LookupCache.CodePages) {
static_assert(sizeof(Page) == 8, "Breaking change in code cache data layout");
::write(fd, &Page, sizeof(Page));
uint64_t NumEntrypoints = Entrypoints.size();
::write(fd, &NumEntrypoints, sizeof(NumEntrypoints));
::write(fd, Entrypoints.data(), Entrypoints.size() * sizeof(Entrypoints[0]));
}
return true;
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
// Generate a literal so we can place it
uint64_t Pointer = ForStorage ? 0 : GetNamedSymbolLiteral(CTX, Reloc.NamedSymbolLiteral.Symbol);
Emitter.dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = ForStorage ? 0 : reinterpret_cast<uint64_t>(CTX.ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Emitter.dc64(GuestEntry + Reloc.GuestRIP.GuestRIP);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
uint64_t Pointer = Reloc.GuestRIP.GuestRIP + GuestEntry;
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIP.RegisterIndex), Pointer, true);
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type {}", ToUnderlying(Reloc.Header.Type));
}
}
return true;
}
+47 -38
View File
@@ -437,6 +437,10 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
Profiler::PostForkAction(Child);
if (Child) {
if (CodeMapWriter) {
CodeMapWriter->ResetAfterFork();
}
CodeInvalidationMutex.StealAndDropActiveLocks();
if (Config.StrictInProcessSplitLocks) {
StrictSplitLockMutex = 0;
@@ -459,9 +463,14 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
}
#endif
void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
void ContextImpl::OnCodeBufferAllocated(const fextl::shared_ptr<CPU::CodeBuffer>& Buffer) {
if (Config.GlobalJITNaming()) {
Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
Symbols.RegisterJITSpace(Buffer->Ptr, Buffer->Size);
}
{
std::scoped_lock lk {CodeBufferListLock};
CodeBufferList.emplace_back(Buffer);
}
}
@@ -715,7 +724,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (SourcecodeResolver && Config.GDBSymbols()) {
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (MappedSection) {
MappedSection->FileInfo.SourcecodeMap = SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, MappedSection->FileInfo.FileId);
MappedSection->FileInfo.SourcecodeMap =
SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, CodeMap::GetBaseFilename(MappedSection->FileInfo, false));
}
}
@@ -833,10 +843,14 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
Thread->CPUBackend->ClearRelocations();
}
fextl::vector<uint64_t> CodePages;
if (NeedsAddGuestCodeRanges) {
// Track in the guest to host map all entrypoints for all pages the compiled block touches, if any page didn't previously
// contain code, inform the frontend so it can setup SMC detection.
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
CodePages.reserve(BlockInfo->CodePages.size());
CodePages.insert(CodePages.end(), BlockInfo->CodePages.begin(), BlockInfo->CodePages.end());
for (auto CodePage : BlockInfo->CodePages) {
if (Thread->LookupCache->AddBlockExecutableRange(Thread, BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
@@ -845,8 +859,16 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
// Insert to lookup cache
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, HostAddr);
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, CodePages, HostAddr);
}
if (CodeMapWriter) {
auto Region = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (Region && Region->FileStartVA != 0) {
CodeMapWriter->AppendBlock(*Region, GuestRIP);
}
}
return (uintptr_t)CodePtr;
@@ -873,50 +895,37 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
void ContextImpl::InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) {
FEXCORE_PROFILE_SCOPED("InvalidateCodeBuffersCodeRange");
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::scoped_lock lk {CodeBufferListLock};
auto it = CodeBufferList.begin();
while (it != CodeBufferList.end()) {
if (auto Strong = it->lock()) {
Strong->LookupCache->InvalidateRange(Start, Length);
it++;
} else {
it = CodeBufferList.erase(it);
}
}
}
void ContextImpl::InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
auto lk = Thread->LookupCache->AcquireWriteLock();
auto& CodePages = Thread->LookupCache->Shared->CodePages;
if (Thread->LookupCache->InvalidateCacheRange(Start, Length)) {
FEXCORE_PROFILE_SCOPED("InvalidateCallRet");
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
Accumulator.emplace_back(std::move(it->second));
}
bool InvalidatedAnyEntries = false;
for (const auto& PageEntries : Accumulator) {
for (const auto& Entry : PageEntries) {
if (ContextImpl::ThreadRemoveCodeEntry(Thread, Entry, lk)) {
InvalidatedAnyEntries = true;
}
}
}
if (InvalidatedAnyEntries) {
// This may cause access violations in the thread on Windows as zeroing is not atomic, this is handled by the frontend
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP,
const FEXCore::LookupCacheWriteLockToken& lk) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP, lk);
}
void ContextImpl::ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
static_cast<ContextImpl*>(Frame->Thread->CTX)->SyscallHandler->InvalidateGuestCodeRange(Frame->Thread, GuestRIP, 1);
}
@@ -5,7 +5,6 @@
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Config/Config.h>
@@ -489,7 +488,7 @@ void Dispatcher::EmitDispatcher() {
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->X86CodeGen.CallbackReturn);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->SignalDelegation->GetThunkCallbackRET());
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, CTX->Config.Is64BitMode ? 16 : 12);
@@ -51,6 +51,10 @@ public:
}
#endif
uint64_t GetExitFunctionLinkerAddress() const {
return ExitFunctionLinkerAddress;
}
SignalDelegatorConfig MakeSignalDelegatorConfig() const;
protected:
@@ -9,7 +9,6 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/LookupCache.h"
#include <array>
@@ -90,11 +89,6 @@ Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
}
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
if (EntryPoint == CTX->X86CodeGen.CallbackReturn) {
return true;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
+1 -2
View File
@@ -50,14 +50,13 @@ DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
auto Dst = GetReg(Node);
uint64_t Mask = ~0ULL;
const auto OpSize = IROp->Size;
if (OpSize == IR::OpSize::i32Bit) {
Mask = 0xFFFF'FFFFULL;
}
LoadConstant(ARMEmitter::Size::i64Bit, Dst, Constant & Mask);
InsertGuestRIPMove(GetReg(Node), Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -11,23 +11,18 @@ $end_info$
#include <FEXCore/Core/Thunks.h>
namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl& CTX, FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op)); break;
return CTX.Dispatcher->GetExitFunctionLinkerAddress();
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum& Sum) {
Relocation MoveABI {};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE};
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
@@ -38,9 +33,9 @@ void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR
}
Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
uint64_t Pointer = GetNamedSymbolLiteral(*CTX, Op);
Arm64JITCore::NamedSymbolLiteralPair Lit {
NamedSymbolLiteralPair Lit {
.Lit = Pointer,
.MoveABI =
{
@@ -48,92 +43,72 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
{
.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit) {
switch (Lit.MoveABI.Header.Type) {
case RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL:
case RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Lit.MoveABI.Header.Offset = GetCursorOffset();
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type for {}", __FUNCTION__);
}
BindOrRestart(&Lit.Loc);
dc64(Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
auto Arm64JITCore::InsertGuestRIPLiteral(uint64_t GuestRIP) -> NamedSymbolLiteralPair {
return {
.Lit = GuestRIP,
.MoveABI =
{
.GuestRIP = {.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL,
},
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
.GuestRIP = GuestRIP},
},
};
}
void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant) {
Relocation MoveABI {};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
MoveABI.GuestRIP.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE};
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
MoveABI.GuestRIP.GuestRIP = Constant;
MoveABI.GuestRIP.RegisterIndex = Reg.Idx();
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, false);
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation> Relocations) {
const auto OrigBase = GetBufferBase();
const auto OrigSize = GetBufferSize();
const auto OrigOffset = GetCursorOffset();
SetBuffer(reinterpret_cast<std::uint8_t*>(Code.data()), Code.size_bytes());
for (auto& Reloc : Relocations) {
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc.NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(Reloc.NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIPMove.RegisterIndex), Pointer, true);
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations(uint64_t GuestBaseAddress) {
// Rebase relocations to library base address
for (auto& Relocation : Relocations) {
switch (Relocation.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE:
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Relocation.GuestRIP.GuestRIP -= GuestBaseAddress;
break;
}
default:;
}
}
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return true;
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations() {
return std::move(Relocations);
}
@@ -138,26 +138,6 @@ DEF_OP(CAS) {
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
eor(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
(void)cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
const auto OpSize = IROp->Size;
@@ -328,8 +328,7 @@ DEF_OP(Thunk) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr));
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
InsertNamedThunkRelocation(ARMEmitter::Reg::r2, Op->ThunkNameHash);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
} else {
+58 -41
View File
@@ -493,7 +493,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
static void DirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = JumpThunkStartAddress / 4 - CallerAddress / 4;
@@ -511,11 +511,12 @@ static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Co
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
static void IndirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
BranchEmit.b(0x8);
// Restore branch +2 instructions to jump to the linker block
BranchEmit.b(0x2);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
@@ -577,16 +578,11 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
if (KnownCallMarkerInst == ExpectedKnownCallMarkerInst) {
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Frame, Record, true); }, lk);
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, true); }, lk);
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, false);
},
lk);
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, false); }, lk);
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
@@ -820,19 +816,28 @@ void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool C
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
const auto PrevNumAllocations = Relocations.size();
this->Entry = Entry;
this->DebugData = DebugData;
this->IR = IR;
RequiresFarARM64Jumps = false;
SSANodeMultiplier = 24;
// Prepare restart via long jump in case branch encoding fails.
// This uses UncheckedLongJump since we don't implement std::longjmp in WoA setups
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(RestartControl.RestartJump))) {
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(ThreadState->RestartJump))) {
case RestartOptions::Control::Incoming:
// Nothing
break;
case RestartOptions::Control::EnableFarARM64Jumps: RequiresFarARM64Jumps = true; break;
default: ERROR_AND_DIE_FMT("Unhandled Arm64 restart condition!");
case RestartOptions::Control::NeedsLargerJITSpace:
// Get rid of the claimed buffer immediately, we can't fit in it at all.
TempAllocator.UnclaimBuffer();
SSANodeMultiplier *= 2;
break;
default: LOGMAN_MSG_A_FMT("Unhandled Arm64 restart condition!");
}
uint32_t SSACount = IR->GetSSACount();
@@ -844,12 +849,19 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CodeData.EntryPoints.clear();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = 0x1000 + SSACount * 24;
// One page baseline, plus SSANodeMultipler bytes, plus another page for guard page.
const uint32_t DesiredBufferRange = AlignUp(FEXCore::Utils::FEX_PAGE_SIZE * 2 + SSACount * SSANodeMultiplier, FEXCore::Utils::FEX_PAGE_SIZE);
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBuffer = TempAllocator.ReownOrClaimBuffer(BufferRange);
SetBuffer(TempCodeBuffer, BufferRange);
auto TempCodeBufferInfo = TempAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBuffer = TempCodeBufferInfo.Ptr;
const uint32_t UsableBufferRange = TempCodeBufferInfo.Size - FEXCore::Utils::FEX_PAGE_SIZE;
SetBuffer(TempCodeBuffer, UsableBufferRange);
ThreadState->JITGuardPage = reinterpret_cast<uintptr_t>(TempCodeBuffer) + UsableBufferRange;
ThreadState->JITGuardOverflowArgument = FEXCore::ToUnderlying(RestartOptions::Control::NeedsLargerJITSpace);
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
@@ -979,22 +991,28 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// This is a ExitFunctionLinkData struct
BindOrRestart(&l_ExitLink);
dc64(0); // HostCode
dc64(PendingJumpThunk.GuestRIP); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
dc64(0); // HostCode
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(PendingJumpThunk.GuestRIP)); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
BindOrRestart(&l_ExitLink);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
PlaceNamedSymbolLiteral(InsertNamedSymbolLiteral(RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER));
// CodeSize not including the header or tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeBegin;
// Add the JitCodeTail
// Add the JitCodeTail (written later)
Align(alignof(JITCodeTail));
auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
const auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
JITCodeTail JITBlockTail {
.RIP = Entry,
.GuestSize = Size,
.SpinLockFutex = 0,
.SingleInst = SingleInst,
};
// Entries that live after the JITCodeTail.
// These entries correlate JIT code regions with guest RIP regions.
@@ -1012,23 +1030,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// FEXCore::Utils::vl64 GuestRIPOffset;
// };
auto JITRIPEntriesBegin = GetCursorAddress<uint8_t*>();
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
JITBlockTail->GuestSize = Size;
JITBlockTail->SingleInst = SingleInst;
JITBlockTail->SpinLockFutex = 0;
const auto JITRIPEntriesBegin = JITBlockTailLocation + sizeof(JITBlockTail);
auto JITRIPEntriesLocation = JITRIPEntriesBegin;
{
// Store the RIP entries.
JITBlockTail->NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail->OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
JITBlockTail.NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail.OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
uintptr_t CurrentRIPOffset = 0;
uint64_t CurrentPCOffset = 0;
@@ -1044,14 +1052,20 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
}
CursorIncrement(JITRIPEntriesLocation - JITRIPEntriesBegin);
SetCursorOffset(JITRIPEntriesLocation - CodeData.BlockBegin);
Align();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
// Finalize and write block tail data
JITBlockTail.Size = CodeData.Size;
{
auto PrevCur = GetCursorOffset();
memcpy(JITBlockTailLocation, &JITBlockTail, sizeof(JITBlockTail));
SetCursorOffset(JITBlockTailLocation - CodeData.BlockBegin + offsetof(JITCodeTail, RIP));
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(JITBlockTail.RIP));
SetCursorOffset(PrevCur);
}
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
@@ -1060,7 +1074,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Query size of generated code
const auto TempSize = GetCursorOffset();
LOGMAN_THROW_A_FMT(TempSize <= BufferRange, "Exceeded bounds of temporary buffer ({:#x} vs {:#x})", TempSize, BufferRange);
// Bring CodeBuffer up to date
{
@@ -1092,6 +1105,10 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
CodeBegin += Delta;
for (std::size_t Idx = PrevNumAllocations; Idx != Relocations.size(); ++Idx) {
Relocations[Idx].Header.Offset += CodeBuffers.LatestOffset;
}
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
+41 -22
View File
@@ -68,10 +68,10 @@ private:
const bool HostSupportsAFP {};
struct RestartOptions {
FEXCore::UncheckedLongJump::JumpBuf RestartJump;
enum class Control : uint64_t {
Incoming = 0,
EnableFarARM64Jumps = 1,
NeedsLargerJITSpace = 2,
};
};
@@ -79,6 +79,8 @@ private:
// In the rare case when those assumptions are broken, FEX needs to safely restart the JIT.
RestartOptions RestartControl {};
bool RequiresFarARM64Jumps {};
// Default to 6 instructions per SSA node.
uint32_t SSANodeMultiplier {24};
ARMEmitter::BiDirectionalLabel* PendingTargetLabel {};
ARMEmitter::BiDirectionalLabel* PendingCallReturnTargetLabel {};
@@ -360,7 +362,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -371,7 +373,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -392,7 +394,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -413,7 +415,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -434,7 +436,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -455,7 +457,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -476,29 +478,37 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adr_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADR.");
}
return;
}
if (adr(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADR currently unsupported!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adrp_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADRP.");
}
return;
}
if (adrp(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADRP currently unsupported!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -513,7 +523,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
@@ -526,8 +536,6 @@ private:
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
@@ -564,19 +572,30 @@ private:
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Inserts a relocation for a constant value relative to the guest entrypoint
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
NamedSymbolLiteralPair InsertGuestRIPLiteral(uint64_t GuestRIP);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit);
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit);
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation>);
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() override;
/**
* Returns any relocations generated since the last call to TakeRelocations.
*
* GuestBaseAddress must match the base virtual address to which the
* input x86 binary is mapped.
*/
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) override;
/** @} */
+34 -24
View File
@@ -1,79 +1,89 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
enum class RelocationTypes : uint8_t {
enum class RelocationTypes : uint32_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// 4 instruction constant generation
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// 8 byte literal (relative to binary base address)
RELOC_GUEST_RIP_LITERAL,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
// 4 instruction constant generation
// Aligned to struct RelocGuestRIP
RELOC_GUEST_RIP_MOVE,
};
struct RelocationTypeHeader final {
struct FEX_PACKED RelocationHeader final {
// Offset to the relocated host code data
uint64_t Offset {};
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
enum class NamedSymbol : uint32_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
RelocationTypeHeader Header {};
RelocationHeader Header {};
NamedSymbol Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
uint32_t Pad[8];
};
struct RelocNamedThunkMove final {
RelocationTypeHeader Header {};
RelocationHeader Header {};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
uint32_t RegisterIndex;
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
struct RelocGuestRIPMove final {
RelocationTypeHeader Header {};
struct RelocGuestRIP final {
RelocationHeader Header {};
// GPR index the constant is being moved to
// GPR index the constant is being moved to (for non-literal relocations)
uint8_t RegisterIndex;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
char Pad[3];
// The unrelocated RIP that is being moved
// The base RIP (to be moved by the register for non-literal relocations).
// In a serialized code cache, this is relative to the binary base address.
uint64_t GuestRIP;
uint32_t pad2[6] {};
};
union Relocation {
RelocationTypeHeader Header {};
RelocationHeader Header {};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIPMove GuestRIPMove;
RelocGuestRIP GuestRIP;
};
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl&, RelocNamedSymbolLiteral::NamedSymbol);
} // namespace FEXCore::CPU
@@ -75,7 +75,7 @@ LookupCache::~LookupCache() {
// These will get freed when their memory allocators are deallocated.
}
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheReadLockToken& lk) {
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer),
@@ -86,12 +86,12 @@ void LookupCache::ClearL2Cache(const FEXCore::LookupCacheReadLockToken& lk) {
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
CachedCodePages.clear();
}
void LookupCache::ClearCache(const LookupCacheWriteLockToken& lk) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
ClearThreadLocalCaches(lk);
Shared->ClearCache(lk);
}
+76 -34
View File
@@ -8,6 +8,7 @@
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/robin_set.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/memory_resource.h>
@@ -17,7 +18,13 @@
#include <mutex>
namespace FEXCore {
struct LookupCacheWriteLockToken {
struct LookupCacheBaseLockToken {
protected:
// Protected constructor - only derived classes can construct
LookupCacheBaseLockToken() = default;
};
struct LookupCacheWriteLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
@@ -26,7 +33,7 @@ private:
std::lock_guard<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct LookupCacheReadLockToken {
struct LookupCacheReadLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
@@ -74,42 +81,61 @@ struct GuestToHostMap {
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
struct BlockEntry {
uint64_t HostCode;
fextl::vector<uint64_t> CodePages;
};
fextl::robin_map<uint64_t, BlockEntry> BlockList;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap();
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode, const LookupCacheWriteLockToken&) {
const BlockEntry& AddBlockMapping(uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode, const LookupCacheWriteLockToken&) {
// This may replace an existing mapping
// NOTE: Generally no previous entry should exist, however there is one exception:
// If the backend updates the active thread's CodeBuffer, the new associated LookupCache
// may already contain the block address. Since is comparatively rare, we'll just leak
// one of the two blocks in this case.
BlockList[Address] = (uintptr_t)HostCode;
return BlockList.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, CodePages}).first->second;
}
std::optional<uintptr_t> FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
auto HostCode = BlockList.find(Address);
if (HostCode == BlockList.end()) {
return std::nullopt;
return nullptr;
}
return HostCode->second;
return &HostCode->second;
}
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken&) {
bool Erase(uint64_t Address, const LookupCacheWriteLockToken&) {
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
it->second(it->first.HostLink);
}
// Remove from BlockList
return BlockList.erase(Address) != 0;
}
void InvalidateRange(uint64_t Start, uint64_t Length) {
auto lk = AcquireWriteLock();
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
Erase(Entry, lk);
}
}
CodePages.erase(lower, upper);
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
const FEXCore::Context::BlockDelinkerFunc& delinker, const LookupCacheWriteLockToken&) {
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
@@ -185,10 +211,10 @@ public:
if (!HostPtr) {
// Try L3
auto HostCode = Shared->FindBlock(Address, lk);
if (HostCode) {
CacheBlockMapping(Address, HostCode.value(), lk);
HostPtr = HostCode.value();
auto Entry = Shared->FindBlock(Address, lk);
if (Entry) {
CacheBlockMapping(Address, *Entry, false, lk);
HostPtr = Entry->HostCode;
}
}
}
@@ -259,32 +285,25 @@ public:
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, void* HostCode) {
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
Shared->AddBlockMapping(Address, HostCode, lk);
const auto& Entry = Shared->AddBlockMapping(Address, CodePages, HostCode, lk);
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
CacheBlockMapping(Address, Entry, true, lk);
}
// NOTE: It's the caller's responsibility to call Erase() for all other
// GuestToHostMaps that share the same LookupCache. Otherwise, the
// L1/L2 caches will contain stale references to deallocated memory.
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken& lk) {
bool ErasedAny = Shared->Erase(Frame, Address, lk);
// Invalidates L1/L2 for a given guest block
void InvalidateCache(uint64_t Address, const LookupCacheWriteLockToken& lk) {
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
ErasedAny = true;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
@@ -300,7 +319,7 @@ public:
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return ErasedAny;
return;
}
// Page exists, just set the offset to zero
@@ -308,7 +327,23 @@ public:
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
}
return true;
}
// Invalidates all L1/L2 entries for all guest block that intersect the given range
bool InvalidateCacheRange(uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireWriteLock();
auto lower = CachedCodePages.lower_bound(Start >> 12);
auto upper = CachedCodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
InvalidateCache(Entry, lk);
}
}
bool ret = upper != lower;
CachedCodePages.erase(lower, upper);
return ret;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
@@ -317,7 +352,7 @@ public:
}
void ClearCache(const LookupCacheWriteLockToken&);
void ClearL2Cache(const LookupCacheReadLockToken&);
void ClearL2Cache(const LookupCacheBaseLockToken&);
void ClearThreadLocalCaches(const LookupCacheWriteLockToken&);
uintptr_t GetL1Pointer() const {
@@ -345,13 +380,17 @@ public:
}
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode, const LookupCacheReadLockToken& lk) {
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheBaseLockToken& lk) {
for (const auto& CodePage : Entry.CodePages) {
CachedCodePages[CodePage >> 12].insert(Address);
}
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = HostCode;
L1Entry.HostCode = Entry.HostCode;
if (!DisableL2Cache()) {
if (!DisableL2Cache() && !L1Only) {
// Do ful map
auto FullAddress = Address;
Address = Address & (VirtualMemSize - 1);
@@ -368,7 +407,7 @@ private:
if (!NewPageBacking) {
// Couldn't allocate, clear L2 and retry
ClearL2Cache(lk);
CacheBlockMapping(Address, HostCode, lk);
CacheBlockMapping(Address, Entry, false, lk);
return;
}
Pointers[Address] = NewPageBacking;
@@ -380,7 +419,7 @@ private:
// This silently replaces existing mappings
BlockPointers[PageOffset].GuestCode = FullAddress;
BlockPointers[PageOffset].HostCode = HostCode;
BlockPointers[PageOffset].HostCode = Entry.HostCode;
}
}
@@ -398,6 +437,9 @@ private:
return PageMemory + NewBase;
}
// Maps from a page index to all blocks in the page that have at some point been fetched into L1/L2
fextl::map<uint64_t, fextl::robin_set<uint64_t>> CachedCodePages;
uintptr_t PagePointer;
uintptr_t PageMemory;
uintptr_t L1Pointer;
@@ -514,18 +514,17 @@ void OpDispatchBuilder::CALLOp(OpcodeArgs) {
BlockSetRIP = true;
// Call instruction only uses up to 32-bit signed displacement
int64_t TargetOffset = Op->Src[0].Literal();
const int64_t TargetOffset = Op->Src[0].Literal();
auto ConstantPC = GetRelocatedPC(Op);
const auto ConstantPC = GetRelocatedPC(Op);
// Push the return address.
Push(GPRSize, ConstantPC);
const uint64_t NextRIP = Op->PC + Op->InstSize;
uint64_t TargetRIP = NextRIP + TargetOffset;
if (NextRIP != TargetRIP) {
if (TargetOffset != 0) {
// Store the RIP
const uint64_t NextRIP = Op->PC + Op->InstSize;
ExitRelocatedPC(Op, TargetOffset, BranchHint::Call, ConstantPC, [&]() {
auto CallReturnJumpTarget = JumpTargets.find(NextRIP);
if (CallReturnJumpTarget != JumpTargets.end() && CallReturnJumpTarget->second.IsEntryPoint) {
@@ -2750,7 +2749,8 @@ void OpDispatchBuilder::NOTOp(OpcodeArgs) {
if (DestIsLockedMem(Op)) {
HandledLock = true;
Ref DestMem = MakeSegmentAddress(Op, Op->Dest);
_AtomicXor(Size, MaskConst, DestMem);
// Result unused
_AtomicFetchXor(Size, MaskConst, DestMem);
} else if (!Op->Dest.IsGPR()) {
// GPR version plays fast and loose with sizes, be safe for memory tho.
Ref Src = LoadSourceGPR(Op, Op->Dest, Op->Flags);
@@ -3199,7 +3199,7 @@ void OpDispatchBuilder::DECOp(OpcodeArgs) {
void OpDispatchBuilder::STOSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("STOSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3244,7 +3244,7 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("MOVSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3297,45 +3297,57 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
_StoreMem(RegClass::GPR, Size, Src, RDI, Invalid(), OpSize::i8Bit, MemOffsetType::SXTX, 1);
}
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
RSI = Add(OpSize::i64Bit, RSI, PtrDir);
RDI = Add(OpSize::i64Bit, RDI, PtrDir);
RSI = OffsetByDir(RSI, IR::OpSizeToSize(Size));
RDI = OffsetByDir(RDI, IR::OpSizeToSize(Size));
StoreGPRRegister(X86State::REG_RSI, RSI);
StoreGPRRegister(X86State::REG_RDI, RDI);
}
}
IR::OpSize OpDispatchBuilder::GetStringOpSize(X86Tables::DecodedOp Op) const {
LOGMAN_THROW_A_FMT(Is64BitMode || !(Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE), "Invalid modifier on 32bit address");
return !Is64BitMode || (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) ? OpSize::i32Bit : OpSize::i64Bit;
}
void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("CMPSOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
bool Repeat = Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX);
if (!Repeat) {
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src2, Src1);
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
Dest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
Dest_RSI = OffsetByDir(Src_RSI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
} else {
// Calculate flags early.
CalculateDeferredFlags();
@@ -3351,7 +3363,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
SetCurrentCodeBlock(BeforeLoop);
StartNewBlock();
ForeachDirection([this, Op, Size, REPE](int32_t PtrDir) {
ForeachDirection([this, Op, Size, AddrSize, REPE](int32_t PtrDir) {
IRPair<IROp_CondJump> InnerJump;
auto JumpIntoLoop = Jump();
@@ -3363,10 +3375,11 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Working loop
{
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPR(Size, Dest_RSI, Size);
@@ -3383,13 +3396,21 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
Dest_RDI = Add(AddrSize, Src_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
Dest_RSI = Add(AddrSize, Src_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
// If TailCounter != 0, compare sources.
// If TailCounter == 0, set ZF iff that would break.
@@ -3428,7 +3449,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("LODSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3510,31 +3531,37 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
}
void OpDispatchBuilder::SCASOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("SCASOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX)) != 0;
if (!Repeat) {
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src1, Src2);
// Offset the pointer
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
StoreGPRRegister(X86State::REG_RDI, OffsetByDir(TailDest_RDI, IR::OpSizeToSize(Size)));
Ref TailDest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
} else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
ForeachDirection([this, Op, Size](int32_t Dir) {
ForeachDirection([this, Op, Size, AddrSize](int32_t Dir) {
bool REPE = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX;
auto JumpStart = Jump();
@@ -3558,7 +3585,8 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Working loop
{
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
@@ -3569,7 +3597,7 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
CalculateDeferredFlags();
Ref TailCounter = LoadGPRRegister(X86State::REG_RCX);
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
Ref Src_RDI_Tail = LoadGPRRegister(X86State::REG_RDI, AddrSize);
// Decrement counter
TailCounter = Sub(OpSize::i64Bit, TailCounter, 1);
@@ -3577,9 +3605,13 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
TailDest_RDI = Add(OpSize::i64Bit, TailDest_RDI, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
Ref TailDest_RDI = Add(AddrSize, Src_RDI_Tail, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
CalculateDeferredFlags();
InternalCondJump = CondJumpNZCV(REPE ? CondClass::EQ : CondClass::NEQ);
@@ -1655,6 +1655,9 @@ private:
return IR::SizeToOpSize(GetSrcSize(Op));
}
[[nodiscard]]
IR::OpSize GetStringOpSize(X86Tables::DecodedOp Op) const;
// Set flag tracking to prepare for an operation that directly writes NZCV.
void HandleNZCVWrite() {
CachedNZCV = nullptr;
@@ -552,7 +552,7 @@ void OpDispatchBuilder::AVXInsertScalarRound(OpcodeArgs) {
const uint64_t Mode = Op->Src[2].Literal();
const auto DstSize = GetGuestVectorLength();
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Dest, Op->Src[0], Mode, true);
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Src[0], Op->Src[1], Mode, true);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -60,15 +60,13 @@ void OpDispatchBuilder::SetX87Top(Ref Value) {
// Float LoaD operation with memory operand
void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
Ref ConvertedData = Data;
// Convert to 80bit float
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
ConvertedData = _F80CVTTo(Data, ReadWidth);
ConvertedData = _F80CVTTo(Data, Width);
}
_PushStack(ConvertedData, Data, ReadWidth);
_PushStack(ConvertedData, Data, Width);
}
// Float LoaD operation with memory operand
@@ -80,7 +78,7 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Data, OpSize::i128Bit);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
@@ -92,7 +90,7 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::i128Bit);
_PushStack(Data, Data, OpSize::f80Bit);
}
void OpDispatchBuilder::FILD(OpcodeArgs) {
@@ -123,11 +121,12 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VLoadTwoGPRs(shifted, upper);
_PushStack(ConvertedData, Invalid(), ReadWidth);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Width == OpSize::i32Bit || Width == OpSize::i64Bit || Width == OpSize::f80Bit, "Invalid store width for FST");
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::f80Bit;
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
@@ -877,8 +876,8 @@ void OpDispatchBuilder::X87FXTRACT(OpcodeArgs) {
_PopStackDestroy();
auto Exp = _F80XTRACT_EXP(Top);
auto Sig = _F80XTRACT_SIG(Top);
_PushStack(Exp, Invalid(), OpSize::f80Bit);
_PushStack(Sig, Invalid(), OpSize::f80Bit);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -59,7 +59,6 @@ void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
// F64 ops
// Float load op with memory operand
void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
// Convert to 64bit float
Ref ConvertedData = Data;
@@ -68,7 +67,7 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
} else if (Width == OpSize::f80Bit) {
ConvertedData = _F80CVT(OpSize::i64Bit, Data);
}
_PushStack(ConvertedData, Data, ReadWidth);
_PushStack(ConvertedData, Data, Width);
}
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
@@ -76,7 +75,7 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Data, OpSize::i64Bit);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
@@ -100,7 +99,7 @@ void OpDispatchBuilder::FILDF64(OpcodeArgs) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
auto ConvertedData = _Float_FromGPR_S(OpSize::i64Bit, ReadWidth == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, Data);
_PushStack(ConvertedData, Invalid(), ReadWidth);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
@@ -397,7 +396,7 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::i64Bit);
_PushStack(Sig, Invalid(), OpSize::i64Bit);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -1,89 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
desc: Guest-side assembly helpers used by the backends
$end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#include <cstring>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
X86GeneratedCode::X86GeneratedCode() {
#ifdef _WIN32
// No need to allocate anything in this config.
#else
// Allocate a page for our emulated guest
CodePtr = AllocateGuestCodeSpace(CODE_SIZE);
constexpr std::array<uint8_t, 2> SignalReturnCode = {
0x0F, 0x3E, // CALLBACKRET FEX Instruction
};
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
memcpy(reinterpret_cast<void*>(CallbackReturn), SignalReturnCode.data(), SignalReturnCode.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
#endif
}
X86GeneratedCode::~X86GeneratedCode() {
#ifndef _WIN32
FEXCore::Allocator::VirtualFree(CodePtr, CODE_SIZE);
#endif
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
#ifndef _WIN32
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
auto Result = FEXCore::Allocator::VirtualAlloc(Size);
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(Result), Size);
return Result;
}
// First 64bit page
constexpr uintptr_t LOCATION_MAX = 0x1'0000'0000;
// 32bit mode
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void* Ptr = ::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED && reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
::munmap(Ptr, Size);
continue;
}
if (Ptr != MAP_FAILED) {
return Ptr;
}
}
// Can't do anything about this
// Here's hoping the application doesn't use signals
return MAP_FAILED;
#else
return nullptr;
#endif
}
} // namespace FEXCore
@@ -1,25 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
$end_info$
*/
#pragma once
#include <stddef.h>
#include <stdint.h>
namespace FEXCore {
class X86GeneratedCode final {
public:
X86GeneratedCode();
~X86GeneratedCode();
uint64_t CallbackReturn {};
private:
void* CodePtr {};
void* AllocateGuestCodeSpace(size_t Size);
};
} // namespace FEXCore
+1 -18
View File
@@ -60,24 +60,7 @@ struct NodeID final {
Value = 0;
}
[[nodiscard]] friend constexpr bool operator==(NodeID, NodeID) noexcept = default;
[[nodiscard]]
friend constexpr bool operator<(NodeID lhs, NodeID rhs) noexcept {
return lhs.Value < rhs.Value;
}
[[nodiscard]]
friend constexpr bool operator>(NodeID lhs, NodeID rhs) noexcept {
return operator<(rhs, lhs);
}
[[nodiscard]]
friend constexpr bool operator<=(NodeID lhs, NodeID rhs) noexcept {
return !operator>(lhs, rhs);
}
[[nodiscard]]
friend constexpr bool operator>=(NodeID lhs, NodeID rhs) noexcept {
return !operator<(lhs, rhs);
}
[[nodiscard]] constexpr auto operator<=>(const NodeID&) const noexcept = default;
friend std::ostream& operator<<(std::ostream& out, NodeID ID) {
out << ID.Value;
-10
View File
@@ -804,16 +804,6 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AtomicXor OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer xor",
"IR layout must match Fetch-variant, otherwise DCE IR optimization breaks!"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i8Bit || Size == FEXCore::IR::OpSize::i16Bit || Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AtomicSwap OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer swap"
+1 -2
View File
@@ -353,8 +353,7 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
AddIndent();
*out << "(%" << IR->GetID(BlockNode) << ") "
<< "CodeBlock ";
*out << "(%" << IR->GetID(BlockNode) << ") " << "CodeBlock ";
*out << "%" << BlockIROp->Begin.ID() << ", ";
*out << "%" << BlockIROp->Last.ID() << std::endl;
@@ -65,7 +65,7 @@ public:
int8_t TopOffset = 0;
FixedSizeStack()
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T()}) {}
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T::Invalid}) {}
void push(const T& Value) {
rotate();
@@ -84,7 +84,7 @@ public:
}
void pop() {
buffer.front() = {StackSlot::INVALID, T()};
buffer.front() = {StackSlot::INVALID, T::Invalid};
rotate(false);
}
@@ -102,7 +102,7 @@ public:
void clear() {
for (auto& Elem : buffer) {
Elem = {StackSlot::UNUSED, T()};
Elem = {StackSlot::UNUSED, T::Invalid};
}
TopOffset = 0;
}
@@ -170,13 +170,8 @@ private:
// Helpers
Ref RotateRight8(uint32_t V, Ref Amount);
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
IREmit->_StoreMemFPR(OpSize::i64Bit, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
auto Upper = IREmit->_VExtractToGPR(OpSize::i128Bit, OpSize::i64Bit, StackNode, 1);
@@ -191,7 +186,24 @@ private:
IREmit->_StoreMemGPR(OpSize::i16Bit, Upper, A.Base, A.Index, OpSize::i64Bit, MemOffsetType::SXTX, A.IndexScale);
}
void Store80BitToMem(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else {
F80SplitStore_Helper(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
}
void StoreStackMem_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(!ReducedPrecisionMode, "Full precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -208,17 +220,7 @@ private:
}
case OpSize::f80Bit: {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else { // 80bit requires split-store
F80SplitStore_Helper(Op, StackNode);
}
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -228,6 +230,8 @@ private:
// Performs a store to memory from a value the stack passed in as StackNode.
// This is the version dealing with the reduced precision case.
void StoreStackMem_Reduced_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(ReducedPrecisionMode, "Reduced precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -244,10 +248,9 @@ private:
break;
}
// 80bit requires split-store
case OpSize::f80Bit: {
StackNode = IREmit->_F80CVTTo(StackNode, OpSize::i64Bit);
F80SplitStore_Helper(Op, StackNode);
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -289,7 +292,7 @@ private:
void Reset();
struct StackMemberInfo {
StackMemberInfo() {}
StackMemberInfo() = delete;
StackMemberInfo(Ref Data)
: StackDataNode(Data) {}
StackMemberInfo(Ref Data, Ref Source, OpSize Size)
@@ -301,6 +304,9 @@ private:
OpSize Size;
Ref Node;
};
static const StackMemberInfo Invalid;
// Tuple is only valid if we have information about the Source of the Stack Data Node.
// In it's valid then OpSize is the original source size and Ref is the original source node.
std::optional<StackMemberData> Source {};
@@ -356,6 +362,8 @@ private:
IRListView* IR = nullptr;
};
inline const X87StackOptimization::StackMemberInfo X87StackOptimization::StackMemberInfo::Invalid {nullptr};
inline void X87StackOptimization::InvalidateCaches() {
InvalidateCachedRegs();
ConstantPool.fill(nullptr);
@@ -725,6 +733,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// The optimization should run per-block
Reset();
IREmit->SetCurrentCodeBlock(BlockNode);
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (!LoweredX87(IROp->Op)) {
continue;
@@ -995,8 +1004,16 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// str w2, [x1]
// or similar. As long as the source size and dest size are one and the same.
// This will avoid any conversions between source and stack element size and conversion back.
if (!SlowPath && Value->Source && Value->Source->Size == Op->StoreSize) {
IREmit->_StoreMemFPR(Op->StoreSize, Value->Source->Node, AddrNode, Offset, Align, OffsetType, OffsetScale);
OpSize StoreSize = Op->StoreSize;
LOGMAN_THROW_A_FMT(Op->StoreSize == OpSize::i32Bit || Op->StoreSize == OpSize::i64Bit || Op->StoreSize == OpSize::f80Bit,
"Invalid store size in x87 store stack mem");
if (!SlowPath && Value->Source && Value->Source->Size == StoreSize) {
Ref SourceValue = Value->Source->Node;
if (Op->StoreSize == OpSize::f80Bit) {
Store80BitToMem(Op, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
} else {
IREmit->_StoreMemFPR(StoreSize, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
break;
}
+30 -2
View File
@@ -1,5 +1,8 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
namespace FEXCore::UncheckedLongJump {
#if defined(_M_ARM_64)
@@ -32,7 +35,7 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value) {
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
// x0 contains the jumpbuffer
ldp x19, x20, [x0, #( 0 * 8)];
@@ -58,6 +61,27 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value)
)" ::
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
// First 12 values are registers [x19,x30].
memcpy(&GPRs[19], &Buffer.Registers[0], sizeof(uint64_t) * 12);
// Next 8 values are [D8,D15]
// Retain upper 64-bits of the register, only modifying lower 64-bits.
for (size_t i = 0; i < 8; ++i) {
memcpy(&FPRs[8 + i], &Buffer.Registers[12 + i], sizeof(uint64_t));
}
// Last value is stack pointer
memcpy(&GPRs[31], &Buffer.Registers[20], sizeof(uint64_t));
// Load the expected value in to X0
GPRs[0] = Value;
// Load the PC with the current LR.
*PC = GPRs[30];
}
#else
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
@@ -86,7 +110,7 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value) {
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
.intel_syntax noprefix;
// rdi contains the jumpbuffer
@@ -115,5 +139,9 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value)
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
LOGMAN_MSG_A_FMT("This is unimplemented on x86-64");
}
#endif
} // namespace FEXCore::UncheckedLongJump
+25 -25
View File
@@ -6,6 +6,9 @@
#include <mutex>
#include <type_traits>
#include <FEXCore/fextl/functional.h>
#include <FEXCore/Utils/EnumUtils.h>
namespace FEXCore::Utils::SpinWaitLock {
/**
* @brief This provides routines to implement implement an "efficient spin-loop" using ARM's WFE and exclusive monitor interfaces.
@@ -125,29 +128,20 @@ static inline uint64_t WFELoadAtomic(uint64_t* Futex) {
return Result;
}
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
return;
}
do {
while (!Pred {}(Result, ComparisonValue)) {
Result = LoadExclusive(Futex);
if (Result == ExpectedValue) {
if (Pred {}(Result, ComparisonValue)) {
return;
}
Result = WFELoadAtomic(Futex);
} while (Result != ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
Result = WFELoadAtomic(Futex);
}
}
template<typename T, typename TT>
static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanoseconds& Timeout) {
@@ -207,19 +201,15 @@ static inline T OneShotWFEBitComparison(T* Futex, T Mask, T Comp) {
}
#else
template<typename T, typename TT>
static inline void Wait(T* Futex, TT ExpectedValue) {
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if (Result == ExpectedValue) {
return;
}
do {
while (!Pred {}(Result, ComparisonValue)) {
Result = AtomicFutex.load();
} while (Result != ExpectedValue);
}
}
template<typename T, typename TT>
@@ -250,6 +240,16 @@ static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanosecon
}
#endif
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
WaitPred<std::equal_to<>, T>(Futex, ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
template<typename T>
static inline void lock(T* Futex) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
+31 -34
View File
@@ -17,7 +17,6 @@
namespace FEXCore::Utils::WritePriorityMutex {
#if !defined(_WIN32)
// A custom mutex that prioritizes exclusive locks.
// In highly contested scenarios, this can help minimize overall contention time.
//
@@ -125,7 +124,6 @@ public:
uint32_t Desired {};
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
@@ -232,8 +230,17 @@ public:
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
#if !defined(_WIN32)
// Initialize the internal mutex object to its default initializer state.
// Should only ever be used in the child process when a Linux fork() has occured.
void StealAndDropActiveLocks() {
Futex = 0;
}
#endif
private:
#if !defined(_WIN32)
void FutexWaitForWriteAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
@@ -255,6 +262,28 @@ private:
// Wake all readers.
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, INT_MAX, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
#else
// Writers wait for the full 32-bit futex.
void FutexWaitForWriteAvailable(uint32_t Expected) {
WaitOnAddress(&Futex, &Expected, sizeof(Futex), INFINITE);
}
// Readers wait for Futex bits [31:16] to be zero.
void FutexWaitForReadAvailable(uint32_t Expected) {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
uint16_t smol_Expected = Expected >> 16;
WaitOnAddress(ReadWaiterAddress, &smol_Expected, sizeof(smol_Expected), INFINITE);
}
void FutexWakeWriter() {
WakeByAddressSingle(&Futex);
}
void FutexWakeReaders() {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
WakeByAddressAll(ReadWaiterAddress);
}
#endif
// Reuse the SpinWaitLock WFE implementations for read/write lock acquiring with WFE.
// Can't reuse the spin-lock directly as some bit-representations are different.
@@ -349,36 +378,4 @@ private:
// Bits[14:0]: Read-owner count.
uint32_t Futex {};
};
#else
// SRWLocks are already write-priority locks in WINE and Windows. Use them to avoid lock stampeding.
class Mutex final {
public:
void lock() {
AcquireSRWLockExclusive(&Futex);
}
void lock_shared() {
AcquireSRWLockShared(&Futex);
}
void unlock() {
ReleaseSRWLockExclusive(&Futex);
}
void unlock_shared() {
ReleaseSRWLockShared(&Futex);
}
bool try_lock() {
return TryAcquireSRWLockExclusive(&Futex);
}
bool try_lock_shared() {
return TryAcquireSRWLockShared(&Futex);
}
private:
SRWLOCK Futex = SRWLOCK_INIT;
};
#endif
} // namespace FEXCore::Utils::WritePriorityMutex
+63 -33
View File
@@ -103,28 +103,25 @@ static inline std::optional<fextl::string> EnumParser(const ArrayPairType& EnumP
return fextl::fmt::format("{}", EnumMask);
}
namespace DefaultValues {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) extern const P(type) P(enum);
#define OPT_STR(group, enum, json, default) extern const std::string_view P(enum);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#include <FEXCore/Config/ConfigValues.inl>
using StringArrayType = fextl::list<fextl::string>;
namespace Type {
using StringArrayType = fextl::list<fextl::string>;
#define OPT_BASE(type, group, enum, json, default) using P(enum) = P(type);
#define OPT_STR(group, enum, json, default) using P(enum) = fextl::string;
#define OPT_STRARRAY(group, enum, json, default) using P(enum) = StringArrayType;
namespace detail {
template<ConfigOption Option>
struct ConfigOptionInfo;
#define DEFINE_METAINFO(type, enum, default) \
template<> \
struct ConfigOptionInfo<ConfigOption::CONFIG_##enum> { \
using Type = type; \
static auto Default() { \
extern default; \
return enum; \
} \
};
#define OPT_BASE(type, group, enum, json, default) DEFINE_METAINFO(type, enum, const type enum)
#define OPT_STR(group, enum, json, default) DEFINE_METAINFO(fextl::string, enum, const std::string_view enum)
#define OPT_STRARRAY(group, enum, json, default) DEFINE_METAINFO(StringArrayType, enum, const std::string_view enum)
#include <FEXCore/Config/ConfigValues.inl>
} // namespace Type
#define FEX_CONFIG_OPT(name, enum) \
FEXCore::Config::Value<FEXCore::Config::DefaultValues::Type::enum> name { \
FEXCore::Config::CONFIG_##enum, \
FEXCore::Config::DefaultValues::enum \
}
#undef P
} // namespace DefaultValues
} // namespace detail
FEX_DEFAULT_VISIBILITY void SetDataDirectory(std::string_view Path, bool Global);
FEX_DEFAULT_VISIBILITY void SetConfigDirectory(const std::string_view Path, bool Global);
@@ -135,8 +132,7 @@ FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigDirectory(bool Global);
FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigFileLocation(bool Global = false);
FEX_DEFAULT_VISIBILITY fextl::string GetApplicationConfig(const std::string_view Program, bool Global);
using LayerValue =
std::variant< fextl::string, DefaultValues::Type::StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerValue = std::variant< fextl::string, StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerOptions = fextl::unordered_map<ConfigOption, LayerValue>;
@@ -151,16 +147,16 @@ public:
return OptionMap.find(Option) != OptionMap.end();
}
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
std::optional<StringArrayType*> All(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
return std::nullopt;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
return &std::get<DefaultValues::Type::StringArrayType>(Value);
return &std::get<StringArrayType>(Value);
}
std::optional<fextl::string*> Get(ConfigOption Option) {
@@ -201,12 +197,12 @@ public:
auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
// If the option didn't exist as a StringArrayType yet, emplace it.
it = OptionMap.emplace(Option, DefaultValues::Type::StringArrayType {}).first;
it = OptionMap.emplace(Option, StringArrayType {}).first;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<DefaultValues::Type::StringArrayType>(Value).emplace_back(Data);
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<StringArrayType>(Value).emplace_back(Data);
}
void Erase(ConfigOption Option) {
@@ -236,7 +232,9 @@ FEX_DEFAULT_VISIBILITY fextl::string FindContainerPrefix();
FEX_DEFAULT_VISIBILITY void AddLayer(fextl::unique_ptr<FEXCore::Config::Layer> _Layer);
FEX_DEFAULT_VISIBILITY bool Exists(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<StringArrayType*> All(ConfigOption Option);
template<typename T>
FEX_DEFAULT_VISIBILITY std::optional<T> GetConv(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<fextl::string*> Get(ConfigOption Option);
FEX_DEFAULT_VISIBILITY void Set(ConfigOption Option, std::string_view Data);
FEX_DEFAULT_VISIBILITY void Erase(ConfigOption Option);
@@ -271,18 +269,18 @@ public:
return ValueData;
}
Value(T Value) requires (!std::is_same_v<T, DefaultValues::Type::StringArrayType>)
Value(T Value) requires (!std::is_same_v<T, StringArrayType>)
{
ValueData = std::move(Value);
}
// Array value types.
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, StringArrayType>)
{
GetListIfExists(Option, &ValueData);
}
DefaultValues::Type::StringArrayType& All() requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
StringArrayType& All() requires (std::is_same_v<T, StringArrayType>)
{
return ValueData;
}
@@ -293,6 +291,38 @@ private:
static T GetIfExists(FEXCore::Config::ConfigOption Option, T Default);
static T GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
};
/**
* Wrapper around Value that automatically picks the default for the given ConfigOption
*/
template<ConfigOption Option>
struct FEX_DEFAULT_VISIBILITY Getter : public Value<typename detail::ConfigOptionInfo<Option>::Type> {
using OptionInfo = detail::ConfigOptionInfo<Option>;
Getter()
: Value<typename OptionInfo::Type> {Option, OptionInfo::Default()} {}
};
/**
* Helper for reading a config value with caching.
*
* Typically this is used to declare class members so that the value is read
* on construction of the parent.
*/
#define FEX_CONFIG_OPT(name, enum) FEXCore::Config::Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> name {}
#define OPT_BASE(type, group, enum, json, default) \
/** \
* Helper for reading a config value. \
* \
* In contrast to FEX_CONFIG_OPT, this can be used in arbitrary expressions, \
* at the expense of not caching the value. Use Getter instead if the value \
* is read frequently. \
*/ \
inline auto Get_##enum() { \
return Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> {}; \
}
#include <FEXCore/Config/ConfigValues.inl>
} // namespace FEXCore::Config
+136 -1
View File
@@ -1,10 +1,20 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/functional.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <atomic>
#include <cstdint>
#include <mutex>
#include <optional>
#include <shared_mutex>
#include <span>
#include <unistd.h>
namespace FEXCore {
@@ -20,8 +30,14 @@ namespace HLE {
struct ExecutableFileInfo {
~ExecutableFileInfo();
#if __clang_major__ < 16
// Workaround for broken aggregate-initialization with std::piecewise_construct
ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap>, uint64_t, fextl::string);
ExecutableFileInfo() = default;
#endif
fextl::unique_ptr<HLE::SourcecodeMap> SourcecodeMap;
fextl::string FileId;
uint64_t FileId = 0;
fextl::string Filename;
};
@@ -33,10 +49,129 @@ struct ExecutableFileSectionInfo {
uintptr_t FileStartVA;
};
using CodeMapFileId = uint64_t;
/**
* Code maps capture information required for offline code cache generation
* and are written to disk during execution of FEX.
*
* Almost all CodeMap data will be an Entry that indicates blocks to be
* compiled for cache generation. The reserved value `LoadExternalLibrary`
* indicates that an instance of ExternalLibraryInfo follows (the entry data
* itself should be skipped in that case).
*/
struct CodeMap {
// Describes the location of an entry block compiled during execution
struct FEX_PACKED Entry {
CodeMapFileId FileId;
uint32_t BlockOffset;
};
// Describes an external library referenced during execution
struct ExternalLibraryInfo {
CodeMapFileId ExternalFileId;
// null-terminated file path; EITHER relative to the main executable OR an absolute path OR starting with a magic identifier:
// - WINE/: Path to Wine/Proton installation
// - WINEPREFIX/: Path to Wine/Proton prefix
// - SLR/: Path to Steam Linux Runtime
// At runtime, FEX will always dump absolute paths
char Path[];
// Followed by padding to a 4 byte boundary
};
// Followed by ExternalLibraryInfo
static constexpr Entry LoadExternalLibrary = {0xffff'ffff'ffff'ffff, 0xffff'ffff};
struct FEX_PACKED SetExecutableFileId {
Entry Marker = {0xffff'ffff'ffff'ffff, 0xffff'fffe};
CodeMapFileId ExecutableFileId;
};
struct ParsedContents {
fextl::string Filename;
fextl::set<uint64_t> Blocks;
bool IsExecutable = false;
};
// Follows scheme fileid[-nomb]
// The nomb ("no multiblock") suffix signifies that the code map is for use without multiblock, only.
static fextl::string GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix);
static fextl::map<CodeMapFileId, ParsedContents> ParseCodeMap(std::ifstream& File);
};
struct CodeMapOpener {
virtual ~CodeMapOpener() = default;
virtual int OpenCodeMapFile() = 0;
};
class CodeMapWriter {
public:
CodeMapWriter(CodeMapOpener&, bool OpenEagerly = false);
~CodeMapWriter();
// Checks if writing is enabled. Calls to this functions may also be interpreted as signals that writes are about to happen
bool IsWriteEnabled(const ExecutableFileSectionInfo&);
void ResetAfterFork() {
if (CodeMapFD.value_or(-1) != -1) {
close(CodeMapFD.value());
CodeMapFD.reset();
}
BufferOffset = 0;
KnownFileIds.clear();
}
bool IsBackingFD(int FD) const {
if (FD == CodeMapFD) {
LogMan::Msg::DFmt("Hiding directory entry for code map FD");
return true;
}
return false;
}
void AppendBlock(const FEXCore::ExecutableFileSectionInfo&, uint64_t Entry);
void AppendLibraryLoad(const FEXCore::ExecutableFileInfo&);
void AppendSetMainExecutable(const FEXCore::ExecutableFileInfo&);
// Thread-safely commit any pending data to disk
void Flush(size_t Offset);
private:
// Queues data into an internal ring buffer.
// Call Flush() to commit the data to disk.
void AppendData(std::span<const std::byte> Data);
// Commit given data range to disk
void Flush(size_t Offset, std::unique_lock<std::shared_mutex>&);
std::shared_mutex Mutex;
fextl::vector<std::byte> Buffer;
std::atomic<size_t> BufferOffset {0};
fextl::set<CodeMapFileId> KnownFileIds;
// std::nullopt: We haven't requested a CodeMapFD yet
// value is -1: We requested a CodeMapFD but FEXServer told us not to write any data
// other values: Code map writing is active
std::optional<int> CodeMapFD;
CodeMapOpener& FileOpener;
};
class AbstractCodeCache {
public:
virtual ~AbstractCodeCache() = default;
/**
* Computes a unique identifier for the referenced binary file to be used for
* generating the code map.
* This identifier is independent of FEX build/runtime configuration and
* stable across FEX updates.
*/
virtual uint64_t ComputeCodeMapId(std::string_view Filename, int FD) = 0;
/**
* Loads a code cache from mapped memory and appends it to the current Core state.
* TODO: Optionally recompiles all contained code blocks at runtime for validation.
+5 -5
View File
@@ -42,9 +42,6 @@ enum OperatingMode {
using CodeRangeInvalidationFn = std::function<void(uint64_t start, uint64_t Length)>;
// Nested vector of guest block entrypoints
using InvalidatedEntryAccumulator = fextl::vector<fextl::vector<uint64_t>>;
using CustomIREntrypointHandler = std::function<void(uintptr_t Entrypoint, IR::IREmitter*)>;
using ExitHandler = std::function<void(Core::InternalThreadState* Thread)>;
@@ -139,10 +136,13 @@ public:
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) = 0;
virtual AbstractCodeCache& GetCodeCache() = 0;
virtual void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter>) = 0;
virtual void FlushAndCloseCodeMap() = 0;
FEX_DEFAULT_VISIBILITY virtual void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(
FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void
InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual void
@@ -67,6 +67,10 @@ public:
return Config;
}
virtual uintptr_t GetThunkCallbackRET() const {
return 0;
}
protected:
SignalDelegatorConfig Config;
};
@@ -4,6 +4,7 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/vector.h>
@@ -118,6 +119,10 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
// The low address of the call-ret stack allocation (not including guard pages)
void* CallRetStackBase {};
uintptr_t JITGuardPage {};
uint64_t JITGuardOverflowArgument {};
FEXCore::UncheckedLongJump::JumpBuf RestartJump;
// BaseFrameState should always be at the end, directly before the interrupt fault page
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState {};
+2 -9
View File
@@ -3,6 +3,7 @@
#include <FEXCore/Utils/EnumOperators.h>
#include <compare>
#include <cstdint>
#include <cstring>
@@ -76,15 +77,7 @@ enum IndexNamedVectorConstant : uint8_t {
struct SHA256Sum final {
uint8_t data[32];
[[nodiscard]]
bool operator<(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) < 0;
}
[[nodiscard]]
bool operator==(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) == 0;
}
[[nodiscard]] auto operator<=>(const SHA256Sum&) const noexcept = default;
};
typedef void ThunkedFunction(void* ArgsRv);
+2 -1
View File
@@ -34,5 +34,6 @@ struct JumpBuf {
#endif
[[nodiscard]] FEX_DEFAULT_VISIBILITY uint64_t SetJump(JumpBuf& Buffer);
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(JumpBuf& Buffer, uint64_t Value);
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(const JumpBuf& Buffer, uint64_t Value);
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC);
} // namespace FEXCore::UncheckedLongJump
+1
View File
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <cstddef>
#include <cstdint>
#ifdef _M_X86_64
@@ -28,4 +28,19 @@ inline fextl::string Trim(fextl::string String, std::string_view TrimTokens = "
return RightTrim(LeftTrim(std::move(String), TrimTokens), TrimTokens);
}
inline fextl::string& ReplaceAllInPlace(fextl::string& Str, std::string_view Token, std::string_view New) {
const auto OriginalTokenSize = Token.size();
const auto NewTokenSize = New.size();
size_t TokenPos {};
auto TokenIter = Str.find(Token, TokenPos);
while (TokenIter != Str.npos) {
Str.replace(TokenIter, OriginalTokenSize, New);
TokenPos += NewTokenSize;
TokenIter = Str.find(Token, TokenPos);
}
return Str;
}
} // namespace FEXCore::StringUtils
@@ -3,6 +3,8 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/list.h>
#include <atomic>
@@ -37,12 +39,6 @@ namespace FEXCore::Utils {
*/
class IntrusivePooledAllocator {
public:
template<typename T>
struct AllocationInfo {
T Ptr;
size_t Size;
};
struct MemoryBuffer;
/**
* @brief Container for tracking the buffers
@@ -403,6 +399,42 @@ private:
const char* Name {};
};
/**
* @brief Thread pool allocator that allocates and frees objects that uses mmap, with a guard page.
*
* The last page of the size provided has the guard.
*/
class PooledAllocatorVirtualWithGuard final : public IntrusivePooledAllocator {
public:
PooledAllocatorVirtualWithGuard() = default;
PooledAllocatorVirtualWithGuard(const char* Name)
: Name {Name} {}
virtual ~PooledAllocatorVirtualWithGuard() {
FreeAllBuffers();
}
private:
void* Alloc(size_t Size) override {
auto Ptr = FEXCore::Allocator::VirtualAlloc(Size);
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
if (Name) {
FEXCore::Allocator::VirtualName(Name, Ptr, Size);
}
return Ptr;
}
void Free(void* Ptr, size_t Size) override {
FEXCore::Allocator::VirtualFree(Ptr, Size);
}
const char* Name {};
};
/**
* @brief Wrapper around the pool allocator for delayed pool reclaiming
*
@@ -460,6 +492,11 @@ public:
UnclaimBuffer();
}
struct AllocationInfo {
Type Ptr;
size_t Size;
};
/**
* @brief Return the owned buffer or allocate another one from the `Allocator`
*
@@ -468,9 +505,9 @@ public:
*
* @param NewSize Optional new size for managed data
*
* @return object of type `Type` allocated within the selected buffer
* @return A usable pointer of type `Type` and the size of the backing store.
*/
Type ReownOrClaimBuffer(std::optional<size_t> NewSize = std::nullopt) {
AllocationInfo ReownOrClaimBufferWithSize(std::optional<size_t> NewSize = std::nullopt) {
// Check if we can cheaply re-own a previous buffer
std::optional Buffer =
IntrusivePooledAllocator::IsClientBufferOwned(ClientOwnedFlag) ? Info : ThreadAllocator.TryToReownBuffer(Info, Size, &ClientOwnedFlag);
@@ -493,7 +530,14 @@ public:
// Leaving this here for future excavation that will definitely occur here
// memset((*Info)->Ptr, 0, Size);
return reinterpret_cast<Type>((*Info)->Ptr);
return {
.Ptr = reinterpret_cast<Type>((*Info)->Ptr),
.Size = (*Info)->Size,
};
}
Type ReownOrClaimBuffer(std::optional<size_t> NewSize = std::nullopt) {
return ReownOrClaimBufferWithSize(NewSize).Ptr;
}
/**
+10
View File
@@ -0,0 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/allocator.h>
#include <tsl/robin_set.h>
namespace fextl {
template<class Key, class Hash = std::hash<Key>, class KeyEqual = std::equal_to<Key>, class Allocator = fextl::FEXAlloc<Key>>
using robin_set = tsl::robin_set<Key, Hash, KeyEqual, Allocator>;
}
+22 -16
View File
@@ -95,14 +95,15 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
CHECK(DisassembleEncoding(1) == 0x10fffffe);
}
{
// Will generate nop + adr.
// Will generate nop + nop + adr.
ForwardLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x1000003e);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x1000003e);
}
{
// Will generate adr.
@@ -115,14 +116,15 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
}
{
// Will generate nop + adr.
// Will generate nop + nop + adr.
BiDirectionalLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x1000003e);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x1000003e);
}
{
@@ -143,12 +145,12 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
}
{
// Will generate nop + adrp.
// Will generate nop + nop + adrp.
ForwardLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, and then aligned to a page.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 2); ++i) {
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 3); ++i) {
nop();
}
@@ -156,11 +158,12 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x9000081e);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x9000081e);
}
{
// Will generate adrp + add.
// Will generate nop + adrp + add.
ForwardLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
@@ -172,8 +175,9 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb000081e);
CHECK(DisassembleEncoding(1) == 0x910013de);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xb000081e);
CHECK(DisassembleEncoding(2) == 0x910013de);
}
@@ -195,12 +199,12 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
}
{
// Will generate nop + adrp.
// Will generate nop + nop + adrp.
BiDirectionalLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, and then aligned to a page.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 2); ++i) {
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 3); ++i) {
nop();
}
@@ -208,11 +212,12 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0x9000081e);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x9000081e);
}
{
// Will generate adrp + add.
// Will generate nop + adrp + add.
BiDirectionalLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
@@ -224,8 +229,9 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xb000081e);
CHECK(DisassembleEncoding(1) == 0x910013de);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xb000081e);
CHECK(DisassembleEncoding(2) == 0x910013de);
}
}
TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: Add/subtract immediate") {
+1 -2
View File
@@ -1,6 +1,5 @@
#!/usr/bin/python3
import xxhash
import hashlib
import sys
import os
import shutil
@@ -188,5 +187,5 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
-3
View File
@@ -1,8 +1,5 @@
#!/usr/bin/python3
import os
import subprocess
import sys
import tempfile
import platform
def ListContainsRequired(Features, RequiredFeatures):
+2 -2
View File
@@ -4,7 +4,7 @@ from clang.cindex import CursorKind
from clang.cindex import TypeKind
from clang.cindex import TranslationUnit
import sys
from dataclasses import dataclass, field
from dataclasses import dataclass
import subprocess
import logging
logger = logging.getLogger()
@@ -548,5 +548,5 @@ def main():
PrintFunctionDecls()
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+3 -3
View File
@@ -9,19 +9,19 @@ for fileid in ~/.fex-emu/aotir/*.path; do
else
args="$args --no-abilocalflags"
fi
if [ "${fileid: -7 : 1}" == "T" ]; then
args="$args --tsoenabled"
else
args="$args --no-tsoenabled"
fi
if [ "${fileid: -8 : 1}" == "S" ]; then
args="$args --smc=full"
else
args="$args --smc=mman"
fi
if [ -f "${fileid%.path}.aotir" ]; then
echo "`basename $fileid` has already been generated"
else
+2 -2
View File
@@ -1,5 +1,5 @@
#!/usr/bin/python3
from dataclasses import dataclass, field
from dataclasses import dataclass
import math
import sys
import logging
@@ -282,5 +282,5 @@ def main():
ExportCommonSyscallDefines()
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+26 -21
View File
@@ -2,7 +2,6 @@
import os
import subprocess
import sys
import tempfile
import re
_Arch = None
@@ -203,10 +202,26 @@ def UpdatePPA():
return DidUpdate
def CheckAndInstallPackageUpdates():
PackagesToInstall = GetPackagesToInstall()
def InstallPackages(PackagesToInstall):
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + PackagesToInstall)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages updated")
else:
print("Packages failed to update")
return DidInstall
def CheckAndInstallPackageUpdates(PackagesToInstall, InstallIfNotFound=False):
for Package in PackagesToInstall[:]:
UpgradableStatus = subprocess.check_output(["apt", "list", "--upgradable", Package]).decode("utf-8")
UpgradableStatus = subprocess.check_output(["apt", "list", "--upgradable", Package], stderr=None).decode("utf-8")
Found = False
for Line in UpgradableStatus.split("\n"):
# If the package exists to be upgraded then it will appear in this list
@@ -221,28 +236,14 @@ def CheckAndInstallPackageUpdates():
if Package in Line and "upgradable" in Line:
Found = True
if Found == False:
if InstallIfNotFound == False and Found == False:
PackagesToInstall.remove(Package)
if len(PackagesToInstall) > 0:
print ("Found updates for packages: {}".format(PackagesToInstall))
print ("This bit may ask for your password")
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + PackagesToInstall)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages updated")
else:
print("Packages failed to update")
return DidInstall
return InstallPackages(PackagesToInstall)
return True
@@ -354,10 +355,14 @@ def main():
if not UpdatePPA():
print ("apt sources failed to update. Not continuing")
ExitWithStatus(-1)
if not CheckAndInstallPackageUpdates():
if not CheckAndInstallPackageUpdates(GetPackagesToInstall()):
print ("apt packages failed to update. Not continuing")
ExitWithStatus(-1)
else:
if not CheckAndInstallPackageUpdates(["software-properties-common"], True):
print ("software-properties-common package failed to update. Not continuing")
ExitWithStatus(-1)
if not InstallPPA():
print ("PPA failed to install. Not continuing")
ExitWithStatus(-1)
+2 -2
View File
@@ -1,6 +1,6 @@
#!/usr/bin/python3
import base64
from dataclasses import dataclass, field
from dataclasses import dataclass
from enum import Flag
import json
import struct
@@ -254,5 +254,5 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+2 -2
View File
@@ -4,7 +4,7 @@ from clang.cindex import CursorKind
from clang.cindex import TypeKind
from clang.cindex import TranslationUnit
import sys
from dataclasses import dataclass, field
from dataclasses import dataclass
import subprocess
import logging
logger = logging.getLogger()
@@ -774,5 +774,5 @@ def main():
return Result
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
-4
View File
@@ -1,13 +1,9 @@
#!/usr/bin/python3
from enum import Flag
import json
import os
import struct
import sys
import glob
from threading import Thread
import subprocess
import time
import multiprocessing
from shutil import which
+1 -1
View File
@@ -76,6 +76,6 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
-1
View File
@@ -1,7 +1,6 @@
#!/usr/bin/python3
import re
import sys
import subprocess
try:
from packaging.version import Version as version_check
except:
-1
View File
@@ -1,7 +1,6 @@
#!/bin/env python3
import sys
import fileinput
import re
# Handles the following formats:
+2 -2
View File
@@ -4,8 +4,8 @@ import os
import sys
import subprocess
# Check if FEX indicates support for AVX
def DoesFEXSupportAVX(mode):
# Check if FEX indicates support for AVX
fex_interpreter_path = os.path.dirname(sys.argv[7]) + "/FEX"
args = list()
@@ -22,8 +22,8 @@ def DoesFEXSupportAVX(mode):
return 'avx' in flags and 'avx2' in flags
return False
# Check if the test itself requires AVX
def TestRequiresAVXSupport():
# Check if the test itself requires AVX
exe_path = sys.argv[len(sys.argv) - 1]
json_path = os.path.dirname(os.path.dirname(exe_path)) + '/requirements/' + os.path.basename(exe_path) + '.json'
-3
View File
@@ -1,6 +1,3 @@
from enum import Flag
import json
import struct
import sys
from json_config_parse import parse_json
-3
View File
@@ -1,6 +1,3 @@
from enum import Flag
import json
import struct
import sys
from json_config_parse import parse_json
-3
View File
@@ -3,9 +3,6 @@
# Save current directory
DIR=$(pwd)
# Get the absolute path to the Scripts directory
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Parse arguments
CHANGED_ONLY=false
TARGET_DIR=""
+1 -2
View File
@@ -1,7 +1,6 @@
#!/usr/bin/python3
import sys
import subprocess
import os.path
from os import path
from shutil import which
@@ -77,4 +76,4 @@ if (is_known_failure):
sys.exit(1)
else:
# Just return the result code if we don't have this test as a known failure
sys.exit(ResultCode);
sys.exit(ResultCode)
+3
View File
@@ -14,10 +14,12 @@
#include <span>
#include <utility>
#include <vector>
#include <unistd.h>
#include <FEXCore/fextl/functional.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/LogManager.h>
namespace fasio {
@@ -368,6 +370,7 @@ std::size_t read(AsyncReadStream& Stream, mutable_buffer Buffers, error& ec) {
auto BytesRead = Stream.read_some(Buffers, ec);
TotalBytesRead += BytesRead;
if (Buffers.FD) {
LOGMAN_THROW_A_FMT(**Buffers.FD != -1, "Receiver requested a file descriptor but none was sent");
(void)Buffers.consume_fd();
}
Buffers += BytesRead;
+4 -2
View File
@@ -160,8 +160,10 @@ private:
struct cmsghdr* cmsg = CMSG_FIRSTHDR(&msg);
if (Buffers.FD &&
(cmsg == nullptr || cmsg->cmsg_len != CMSG_LEN(sizeof(int)) || cmsg->cmsg_level != SOL_SOCKET || cmsg->cmsg_type != SCM_RIGHTS)) {
ec = error::invalid;
return 0;
// Not a failure since some data was read for the main message
**Buffers.FD = -1;
ec = error::success;
return BytesRead;
}
if (Buffers.FD) {
+34 -2
View File
@@ -79,8 +79,8 @@ static char* SaveLayerToJSON(char* JsonBuffer, const FEXCore::Config::Layer* Lay
}
if (std::holds_alternative<fextl::string>(it.second)) {
JsonBuffer = json_str(JsonBuffer, Name.data(), std::get<fextl::string>(it.second).c_str());
} else if (std::holds_alternative<FEXCore::Config::DefaultValues::Type::StringArrayType>(it.second)) {
for (auto& var : std::get<FEXCore::Config::DefaultValues::Type::StringArrayType>(it.second)) {
} else if (std::holds_alternative<FEXCore::Config::StringArrayType>(it.second)) {
for (auto& var : std::get<FEXCore::Config::StringArrayType>(it.second)) {
JsonBuffer = json_str(JsonBuffer, Name.data(), var.c_str());
}
} else {
@@ -542,6 +542,13 @@ const char* GetHomeDirectory() {
#endif
fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* DataOverride = getenv("FEX_APP_DATA_LOCATION");
if (PortableInfo.IsPortable && (Global || !DataOverride)) {
@@ -566,6 +573,13 @@ fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableI
}
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (PortableInfo.IsPortable && (Global || !ConfigOverride)) {
return fextl::fmt::format("{}/fex-emu/", PortableInfo.InterpreterPath);
@@ -602,6 +616,24 @@ fextl::string GetConfigDirectory(bool Global, const PortableInformation& Portabl
return ConfigDir;
}
fextl::string GetCacheDirectory() {
const char* CacheOverride = getenv("FEX_APP_CACHE_LOCATION");
if (CacheOverride) {
return CacheOverride;
}
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_SHADER_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* HomeDir = GetHomeDirectory();
const char* CacheXDG = getenv("XDG_CACHE_HOME");
return (CacheXDG ? fextl::string {CacheXDG} : (fextl::string {HomeDir} + "/.cache")) + "/fex-emu/";
}
fextl::string GetConfigFileLocation(bool Global, const PortableInformation& PortableInfo) {
return GetConfigDirectory(Global, PortableInfo) + "Config.json";
}
+1
View File
@@ -62,6 +62,7 @@ const char* GetHomeDirectory();
fextl::string GetDataDirectory(const PortableInformation& PortableInfo);
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo);
fextl::string GetConfigFileLocation(bool Global, const PortableInformation& PortableInfo);
fextl::string GetCacheDirectory();
void InitializeConfigs(const PortableInformation& PortableInfo);
+166 -104
View File
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: MIT
#include "Common/AsyncNet.h"
#include "Common/Config.h"
#include "FDUtils.h"
#include "Common/FEXServerClient.h"
#include <FEXCore/Utils/CompilerDefs.h>
@@ -59,8 +60,8 @@ int RequestPIDFDPacket(int ServerSocket, PacketType Type) {
fasio::mutable_buffer ResBuffer {std::as_writable_bytes(std::span {&Res, 1})};
int NewFD = -1;
ResBuffer.FD = &NewFD;
read(Socket, ResBuffer, ec);
if (ec != fasio::error::success || Res.Header.Type != PacketType::TYPE_SUCCESS) {
auto BytesRead = Socket.read_some(ResBuffer, ec);
if (ec != fasio::error::success || BytesRead != sizeof(Res) || Res.Header.Type != PacketType::TYPE_SUCCESS) {
return -1;
}
@@ -137,15 +138,21 @@ fextl::string GetServerSocketName() {
}
fextl::string GetServerSocketPath() {
fextl::string name {};
#ifndef FEX_STEAM_SUPPORT
FEX_CONFIG_OPT(ServerSocketPath, SERVERSOCKETPATH);
auto name = ServerSocketPath();
name = ServerSocketPath();
if (name.starts_with("/")) {
return name;
}
auto Folder = GetTempFolder();
#else
// Under Steam the FEXServer's socket is a game-specific directory.
auto Folder = GetServerLockFolder();
#endif
if (name.empty()) {
return fextl::fmt::format("{}/{}.FEXServer.Socket", Folder, ::getuid());
@@ -159,25 +166,31 @@ int GetServerFD() {
}
int ConnectToServer(ConnectionOption ConnectionOption) {
auto ServerSocketName = GetServerSocketName();
int SocketFD {-1};
size_t SizeOfAddr {};
struct sockaddr_un addr {};
size_t SizeOfSocketString {};
// Create the initial unix socket
int SocketFD = socket(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0);
SocketFD = socket(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0);
if (SocketFD == -1) {
LogMan::Msg::EFmt("Couldn't open AF_UNIX socket {}", errno);
return -1;
}
// Steam doesn't get to connect to global sockets.
#ifndef FEX_STEAM_SUPPORT
auto ServerSocketName = GetServerSocketName();
// AF_UNIX has a special feature for named socket paths.
// If the name of the socket begins with `\0` then it is an "abstract" socket address.
// The entirety of the name is used as a path to a socket that doesn't have any filesystem backing.
struct sockaddr_un addr {};
addr.sun_family = AF_UNIX;
size_t SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
addr.sun_path[0] = 0; // Abstract AF_UNIX sockets start with \0
strncpy(addr.sun_path + 1, ServerSocketName.data(), SizeOfSocketString);
// Include final null character.
size_t SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
if (connect(SocketFD, reinterpret_cast<struct sockaddr*>(&addr), SizeOfAddr) == -1) {
if (ConnectionOption == ConnectionOption::Default || errno != ECONNREFUSED) {
@@ -186,11 +199,13 @@ int ConnectToServer(ConnectionOption ConnectionOption) {
} else {
return SocketFD;
}
#endif
// Try again with a path-based socket, since abstract sockets will fail if we have been
// placed in a new netns as part of a sandbox.
auto ServerSocketPath = GetServerSocketPath();
addr.sun_family = AF_UNIX;
SizeOfSocketString = std::min(ServerSocketPath.size(), sizeof(addr.sun_path) - 1);
strncpy(addr.sun_path, ServerSocketPath.data(), SizeOfSocketString);
SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
@@ -224,110 +239,125 @@ bool SetupClient(std::string_view InterpreterPath) {
return true;
}
int ConnectToAndStartServer(std::string_view InterpreterPath) {
int ServerFD = ConnectToServer(ConnectionOption::NoPrintConnectionError);
if (ServerFD == -1) {
// Couldn't connect to the server. Start one
int StartServer(std::string_view InterpreterPath, int watch_fd) {
int LocalServerFD {-1};
// Couldn't connect to the server. Start one
// Open some pipes for letting us know when the server is ready
int fds[2] {};
if (pipe2(fds, 0) != 0) {
LogMan::Msg::EFmt("Couldn't open pipe");
// Open some pipes for letting us know when the server is ready
int fds[2] {};
if (pipe2(fds, 0) != 0) {
LogMan::Msg::EFmt("Couldn't open pipe");
return -1;
}
// Extract directory from InterpreterPath
fextl::string InterpreterDir {InterpreterPath};
size_t LastSlash = InterpreterDir.rfind('/');
if (LastSlash != fextl::string::npos) {
InterpreterDir = InterpreterDir.substr(0, LastSlash);
}
fextl::string FEXServerPath = fextl::fmt::format("{}/FEXServer", InterpreterDir);
// Check if a local FEXServer next to FEX exists
// If it does then it takes priority over the installed one
if (!FHU::Filesystem::Exists(FEXServerPath)) {
FEXServerPath = "FEXServer";
}
// Set-up our SIGCHLD handler to ignore the signal.
// This is early in the initialization stage so no handlers have been installed.
//
// We want to ignore the signal so that if FEXServer starts in daemon mode, it
// doesn't leave a zombie process around waiting for something to get the result.
struct sigaction action {};
action.sa_handler = SIG_IGN;
sigaction(SIGCHLD, &action, &action);
pid_t pid = fork();
if (pid == 0) {
// Child
close(fds[0]); // Close read end of pipe
const char* argv[6];
auto pipe_string = fextl::fmt::format("{}", fds[1]);
auto watch_fd_string = fextl::fmt::format("{}", watch_fd);
size_t arg_count {};
argv[arg_count++] = FEXServerPath.c_str();
argv[arg_count++] = "--wait_pipe";
argv[arg_count++] = pipe_string.c_str();
if (watch_fd != -1) {
argv[arg_count++] = "--watch_fd";
argv[arg_count++] = watch_fd_string.c_str();
}
argv[arg_count++] = nullptr;
if (execvp(argv[0], (char* const*)argv) == -1) {
// Let the parent know that we couldn't execute for some reason
uint64_t error {1};
write(fds[1], &error, sizeof(error));
// Give a hopefully helpful error message for users
LogMan::Msg::EFmt("Couldn't execute: {}", argv[0]);
LogMan::Msg::EFmt("This means the squashFS rootfs won't be mounted.");
LogMan::Msg::EFmt("Expect errors!");
// Destroy this fork
exit(1);
}
FEX_UNREACHABLE;
} else {
// Parent
// Wait for the child to exit so we can check if it is mounted or not
close(fds[1]); // Close write end of the pipe
// Wait for a message from FEXServer
pollfd PollFD;
PollFD.fd = fds[0];
PollFD.events = POLLIN | POLLOUT | POLLRDHUP | POLLERR | POLLHUP | POLLNVAL;
// Wait for a result on the pipe that isn't EINTR
while (poll(&PollFD, 1, -1) == -1 && errno == EINTR)
;
// Check if child signaled an error
uint64_t error = 0;
ssize_t bytes_read = read(fds[0], &error, sizeof(error));
close(fds[0]);
if (bytes_read > 0 && error != 0) {
return -1;
}
// Extract directory from InterpreterPath
fextl::string InterpreterDir {InterpreterPath};
size_t LastSlash = InterpreterDir.rfind('/');
if (LastSlash != fextl::string::npos) {
InterpreterDir = InterpreterDir.substr(0, LastSlash);
for (size_t i = 0; i < 5; ++i) {
LocalServerFD = ConnectToServer(ConnectionOption::Default);
if (LocalServerFD != -1) {
break;
}
std::this_thread::sleep_for(std::chrono::seconds(1));
}
fextl::string FEXServerPath = fextl::fmt::format("{}/FEXServer", InterpreterDir);
// Check if a local FEXServer next to FEX exists
// If it does then it takes priority over the installed one
if (!FHU::Filesystem::Exists(FEXServerPath)) {
FEXServerPath = "FEXServer";
if (LocalServerFD == -1) {
// Still couldn't connect to the socket.
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket after launching the process");
}
// Set-up our SIGCHLD handler to ignore the signal.
// This is early in the initialization stage so no handlers have been installed.
//
// We want to ignore the signal so that if FEXServer starts in daemon mode, it
// doesn't leave a zombie process around waiting for something to get the result.
struct sigaction action {};
action.sa_handler = SIG_IGN;
sigaction(SIGCHLD, &action, &action);
pid_t pid = fork();
if (pid == 0) {
// Child
close(fds[0]); // Close read end of pipe
const char* argv[4];
auto pipe_string = fextl::fmt::format("{}", fds[1]);
argv[0] = FEXServerPath.c_str();
argv[1] = "--wait_pipe";
argv[2] = pipe_string.c_str();
argv[3] = nullptr;
if (execvp(argv[0], (char* const*)argv) == -1) {
// Let the parent know that we couldn't execute for some reason
uint64_t error {1};
write(fds[1], &error, sizeof(error));
// Give a hopefully helpful error message for users
LogMan::Msg::EFmt("Couldn't execute: {}", argv[0]);
LogMan::Msg::EFmt("This means the squashFS rootfs won't be mounted.");
LogMan::Msg::EFmt("Expect errors!");
// Destroy this fork
exit(1);
}
FEX_UNREACHABLE;
} else {
// Parent
// Wait for the child to exit so we can check if it is mounted or not
close(fds[1]); // Close write end of the pipe
// Wait for a message from FEXServer
pollfd PollFD;
PollFD.fd = fds[0];
PollFD.events = POLLIN | POLLOUT | POLLRDHUP | POLLERR | POLLHUP | POLLNVAL;
// Wait for a result on the pipe that isn't EINTR
while (poll(&PollFD, 1, -1) == -1 && errno == EINTR)
;
// Check if child signaled an error
uint64_t error = 0;
ssize_t bytes_read = read(fds[0], &error, sizeof(error));
close(fds[0]);
if (bytes_read > 0 && error != 0) {
return -1;
}
for (size_t i = 0; i < 5; ++i) {
ServerFD = ConnectToServer(ConnectionOption::Default);
if (ServerFD != -1) {
break;
}
std::this_thread::sleep_for(std::chrono::seconds(1));
}
if (ServerFD == -1) {
// Still couldn't connect to the socket.
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket {} after launching the process", GetServerSocketName());
}
}
// Restore the original SIGCHLD handler if it existed.
sigaction(SIGCHLD, &action, nullptr);
}
return ServerFD;
// Restore the original SIGCHLD handler if it existed.
sigaction(SIGCHLD, &action, nullptr);
return LocalServerFD;
}
int ConnectToAndStartServer(std::string_view InterpreterPath) {
int LocalServerFD = ConnectToServer(ConnectionOption::NoPrintConnectionError);
if (LocalServerFD == -1) {
LocalServerFD = StartServer(InterpreterPath);
}
return LocalServerFD;
}
/**
@@ -375,6 +405,38 @@ int RequestPIDFD(int ServerSocket) {
return RequestPIDFDPacket(ServerSocket, PacketType::TYPE_GET_PID_FD);
}
int RequestCodeMapFD(int ServerSocket, int ProgramFD, bool HasMultiblock) {
fasio::tcp_socket Socket {ServerSocket};
FEXServerRequestPacket Req {
.Header {
.Type = HasMultiblock ? PacketType::TYPE_QUERY_CODE_MAP : PacketType::TYPE_QUERY_CODE_MAP_NO_MULTIBLOCK,
},
};
// Send request
fasio::error ec;
{
fasio::mutable_buffer WriteBuffer {std::as_writable_bytes(std::span {&Req, 1})};
WriteBuffer.FD = &ProgramFD;
write(Socket, WriteBuffer, ec);
if (ec != fasio::error::success) {
return -1;
}
}
// Wait for success response and log FD
FEXServerResultPacket Res {};
fasio::mutable_buffer ResBuffer {std::as_writable_bytes(std::span {&Res, 1})};
int NewFD = -1;
ResBuffer.FD = &NewFD;
read(Socket, ResBuffer, ec);
if (ec != fasio::error::success || Res.Header.Type != PacketType::TYPE_SUCCESS) {
return -1;
}
return NewFD;
}
/** @} */
/**
+19
View File
@@ -19,6 +19,8 @@ enum class PacketType {
TYPE_GET_LOG_FD,
TYPE_GET_ROOTFS_PATH,
TYPE_GET_PID_FD,
TYPE_QUERY_CODE_MAP,
TYPE_QUERY_CODE_MAP_NO_MULTIBLOCK,
// Result only
TYPE_SUCCESS,
@@ -65,6 +67,13 @@ int GetServerFD();
bool SetupClient(std::string_view InterpreterPath);
/**
* @brief Start a FEXServer instance if possible
*
* @return socket FD for communicating with server
*/
int StartServer(std::string_view InterpreterPath, int watch_fd = -1);
/**
* @brief Connect to and start a FEXServer instance if required
*
@@ -113,6 +122,16 @@ fextl::string RequestRootFSPath(int ServerSocket);
*/
int RequestPIDFD(int ServerSocket);
/**
* @brief Request FEXServer to create a new code map for disk cache population
*
* @param ServerSocket - Socket to the server
* @param ProgramFD - FD for program binary
*
* @return FD to write code map to
*/
int RequestCodeMapFD(int ServerSocket, int ProgramFD, bool HasMultiblock);
/** @} */
/**
+9 -5
View File
@@ -2,9 +2,9 @@
#pragma once
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <optional>
#include <span>
#include <elf.h>
@@ -15,12 +15,16 @@ namespace FEXCore {
* Infers the base virtual address from a file mapping (as described by parameters to a single
* call to mmap()).
*
* Usually the base address can uniquely be inferred, but in edge cases multiple possible
* candidates are returned.
*
* The file offset of any given mapping need not match its virtual address offset from the base
* mapping (file offset = 0). Instead, this function searches the corresponding ELF program headers
* for an entry that generated the given file mapping.
*/
inline std::optional<uint64_t>
inline fextl::vector<uint64_t>
InferMappingBaseAddress(std::span<const Elf64_Phdr> ProgramHeaders, uint64_t Addr, uint64_t Size, uint64_t FileOffset, int AccessFlags) {
fextl::vector<uint64_t> Ret;
for (auto& phdr : ProgramHeaders) {
if (phdr.p_type != PT_LOAD) {
// Skip headers that don't trigger memory mappings
@@ -36,11 +40,11 @@ InferMappingBaseAddress(std::span<const Elf64_Phdr> ProgramHeaders, uint64_t Add
if (FileOffset >= SegmentStartOffset && FileOffset < SegmentStartOffset + phdr.p_filesz &&
(FileOffset & Utils::FEX_PAGE_MASK) == (phdr.p_offset & Utils::FEX_PAGE_MASK)) {
// Compute VA offset relative to the base mapping
return Addr - (phdr.p_vaddr - (phdr.p_offset & 0xfff)) + (ProgramHeaders[0].p_vaddr - (ProgramHeaders[0].p_offset & 0xfff)) -
(FileOffset - SegmentStartOffset);
Ret.push_back(Addr - (phdr.p_vaddr - (phdr.p_offset & 0xfff)) + (ProgramHeaders[0].p_vaddr - (ProgramHeaders[0].p_offset & 0xfff)) -
(FileOffset - SegmentStartOffset));
}
}
return std::nullopt;
return Ret;
}
} // namespace FEXCore
+193 -69
View File
@@ -7,6 +7,9 @@
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/StringUtils.h>
#include <range/v3/view/split.hpp>
#include <range/v3/view/transform.hpp>
#ifdef _M_X86_64
#include "Common/X86Features.h"
#endif
@@ -35,6 +38,22 @@ void FillMIDRInformationViaLinux(FEXCore::HostFeatures* Features) {
#endif
}
#if defined(_M_ARM_64) && !defined(VIXL_SIMULATOR)
__attribute__((naked)) static uint64_t ReadSVEVectorLengthInBits() {
///< Can't use rdvl instruction directly because compilers will complain that sve/sme is required.
__asm(R"(
.word 0x04bf5100 // rdvl x0, #8
ret;
)");
}
#else
[[maybe_unused]]
static int ReadSVEVectorLengthInBits() {
// Return unsupported
return 0;
}
#endif
#ifdef _M_ARM_64
#define GetSysReg(name, reg) \
static uint64_t Get_##name() { \
@@ -53,6 +72,7 @@ GetSysReg(MMFR2_EL1, ID_AA64MMFR2_EL1);
GetSysReg(ZFR0_EL1, s3_0_c0_c4_4); // Can't request by name
GetSysReg(MMFR1_EL1, ID_AA64MMFR1_EL1);
GetSysReg(ISAR2_EL1, ID_AA64ISAR2_EL1);
GetSysReg(DCZID_EL0, DCZID_EL0);
class CPUFeaturesFromID final : public FEX::CPUFeatures {
public:
@@ -66,12 +86,17 @@ public:
MMFR2.SetReg(Get_MMFR2_EL1());
MMFR1.SetReg(Get_MMFR1_EL1());
ISAR2.SetReg(Get_ISAR2_EL1());
DCZID.SetReg(Get_DCZID_EL0());
if (PFR0.SupportsSVE()) {
// Can only query if SVE is supported.
ZFR0.SetReg(Get_ZFR0_EL1());
}
FillFeatureFlags();
if (Supports(CPUFeatures::Feature::SVE2)) {
SVEVL.SetReg(ReadSVEVectorLengthInBits());
}
}
};
@@ -80,6 +105,78 @@ FEX::CPUFeatures GetCPUFeaturesFromIDRegisters() {
}
#endif
class CPUFeaturesFromConfig final : public FEX::CPUFeatures {
public:
CPUFeaturesFromConfig(std::string_view Config) {
auto to_string_view = [](auto rng) {
return std::string_view(&*rng.begin(), ranges::distance(rng));
};
for (auto Option : ranges::views::split(Config, ',') | ranges::views::transform(to_string_view)) {
auto OptionData = ranges::views::split(Option, '=') | ranges::views::transform(to_string_view);
auto OptionDataBegin = ranges::begin(OptionData);
auto OptionDataEnd = ranges::end(OptionData);
if (OptionDataBegin == OptionDataEnd) {
continue;
}
auto Key = *OptionDataBegin;
if (Key.empty()) {
continue;
}
++OptionDataBegin;
if (OptionDataBegin == OptionDataEnd) {
continue;
}
auto Value = *OptionDataBegin;
uint64_t ValueHex {};
char* str_end {};
ValueHex = std::strtoull(Value.data(), &str_end, 16);
if (str_end == Value.data()) {
LogMan::Msg::EFmt("Couldn't parse '{}={}'\n", Key, Value);
continue;
}
if (Key == "isar0") {
ISAR0.SetReg(ValueHex);
} else if (Key == "isar1") {
ISAR1.SetReg(ValueHex);
} else if (Key == "isar2") {
ISAR2.SetReg(ValueHex);
} else if (Key == "pfr0") {
PFR0.SetReg(ValueHex);
} else if (Key == "pfr1") {
PFR1.SetReg(ValueHex);
} else if (Key == "midr") {
MIDR.SetReg(ValueHex);
} else if (Key == "mmfr0") {
MMFR0.SetReg(ValueHex);
} else if (Key == "mmfr1") {
MMFR1.SetReg(ValueHex);
} else if (Key == "mmfr2") {
MMFR2.SetReg(ValueHex);
} else if (Key == "zfr0") {
ZFR0.SetReg(ValueHex);
} else if (Key == "dczid") {
DCZID.SetReg(ValueHex);
} else if (Key == "svevl") {
SVEVL.SetReg(ValueHex);
} else {
LogMan::Msg::EFmt("Unknown Key: {}", Key);
}
}
FillFeatureFlags();
}
};
FEX::CPUFeatures GetCPUFeaturesFromConfig(std::string_view Config) {
return CPUFeaturesFromConfig {Config};
}
class CPUFeaturesAll final : public FEX::CPUFeatures {
public:
CPUFeaturesAll() {
@@ -87,6 +184,9 @@ public:
for (uint32_t i = 0; i < FEXCore::ToUnderlying(FEX::CPUFeatures::Feature::MAX); ++i) {
SetFeature(FEX::CPUFeatures::Feature {i});
}
// Report unsupported for DCZVA
DCZID.SetReg(0b1'0000);
}
};
@@ -346,21 +446,7 @@ void FEX::CPUFeatures::FillFeatureFlags() {
}
}
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
[[maybe_unused]] constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
[[maybe_unused]] constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
[[maybe_unused]]
static uint32_t GetDCZID() {
uint64_t Result {};
__asm("mrs %[Res], DCZID_EL0" : [Res] "=r"(Result));
return Result;
}
static uint32_t GetFPCR() {
uint64_t Result {};
__asm("mrs %[Res], FPCR" : [Res] "=r"(Result));
@@ -371,27 +457,6 @@ static void SetFPCR(uint64_t Value) {
__asm("msr FPCR, %[Value]" ::[Value] "r"(Value));
}
#ifndef VIXL_SIMULATOR
__attribute__((naked)) static uint64_t ReadSVEVectorLengthInBits() {
///< Can't use rdvl instruction directly because compilers will complain that sve/sme is required.
__asm(R"(
.word 0x04bf5100 // rdvl x0, #8
ret;
)");
}
#endif
#else
[[maybe_unused]]
static uint32_t GetDCZID() {
// Return unsupported
return DCZID_DZP_MASK;
}
[[maybe_unused]]
static int ReadSVEVectorLengthInBits() {
// Return unsupported
return 0;
}
#endif
static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEWidth) {
@@ -460,9 +525,76 @@ static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEW
Features->SupportsSVE256 = ForceSVEWidth && ForceSVEWidth >= 256;
}
FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool SupportsCacheMaintenanceOps, uint64_t CTR, uint64_t MIDR) {
FEXCore::HostFeatures HostFeatures;
static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
constexpr uint32_t Implementer_ARM = 0x41;
constexpr uint32_t PartNum_V2 = 0xd4f;
constexpr uint32_t PartNum_V3 = 0xd84;
constexpr uint32_t PartNum_V3AE = 0xd83;
constexpr uint32_t PartNum_X3 = 0xd4e;
constexpr uint32_t PartNum_X4 = 0xd82;
constexpr uint32_t PartNum_X925 = 0xd85;
constexpr uint32_t PartNum_C1Ultra = 0xd8c;
constexpr uint32_t PartNum_C1Premium = 0xd90;
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
auto GetMIDRImplementer = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 24) & 0xFF;
};
auto GetMIDRPartNum = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 4) & 0xFFF;
};
const uint32_t MIDR_Implementer = GetMIDRImplementer(MIDR);
const uint32_t MIDR_PartNum = GetMIDRPartNum(MIDR);
#ifdef _M_ARM_64
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
// - The RNDRRS register will never read a random number. (Always return failure)
// This is contrary to x86 RNG behaviour where it allows spurious failure with RDSEED, but guarantees eventual success.
// This manifested itself on Linux when an x86 processor failed to guarantee forward progress and boot of services would infinite
// loop. Just disable this extension if this CPU is detected.
HostFeatures->SupportsRAND = false;
}
#endif
// The LDAPUR instruction suffers from significant performance issues on many ARM implementations. This is
// listed in the official Cortex errata list as follows:
//
// 3877900
// LDAPUR, LDAPURB, LDAPURH instructions have stricter memory ordering than required
//
// LDAPUR instructions execute with full Load-Acquire ordering instead of the relaxed ordering described
// in the LDAPUR pseudocode. This might cause significant performance degradation in workloads that do
// not require this stricter memory ordering. Note that this erratum only affects the unscaled versions of
// LDAPUR (LDAPUR, LDAPURB, LDAPURH), and not LDAPR (LDAPR, LDAPRB, LDAPRH).
//
// The list of cores to disable its use on was taken from the following LLVM PR that accomplishes the same
// thing: https://github.com/llvm/llvm-project/pull/124274
for (uint32_t CoreIndex = 0; CoreIndex < HostFeatures->CPUMIDRs.size(); CoreIndex++) {
const uint32_t CoreMIDR = HostFeatures->CPUMIDRs[CoreIndex];
const uint32_t Core_MIDR_Implementer = GetMIDRImplementer(CoreMIDR);
const uint32_t Core_MIDR_PartNum = GetMIDRPartNum(CoreMIDR);
bool IgnoreLRCPC2 = (Core_MIDR_Implementer == Implementer_ARM) &&
((Core_MIDR_PartNum == PartNum_V2) || (Core_MIDR_PartNum == PartNum_V3) || (Core_MIDR_PartNum == PartNum_X3) ||
(Core_MIDR_PartNum == PartNum_X4) || (Core_MIDR_PartNum == PartNum_X925) || (Core_MIDR_PartNum == PartNum_V3AE) ||
(Core_MIDR_PartNum == PartNum_C1Ultra) || (Core_MIDR_PartNum == PartNum_C1Premium));
if (IgnoreLRCPC2) {
HostFeatures->SupportsTSOImm9 = false;
break;
}
}
}
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
uint64_t MIDR) {
FEX_CONFIG_OPT(ForceSVEWidth, FORCESVEWIDTH);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
@@ -495,7 +627,7 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
HostFeatures.SupportsSVE256 = ForceSVEWidth() ? ForceSVEWidth() >= 256 : true;
#else
HostFeatures.SupportsSVE128 = Features.Supports(CPUFeatures::Feature::SVE2);
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && ReadSVEVectorLengthInBits() >= 256;
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && Features.GetSVEVectorLengthInBits() >= 256;
#endif
HostFeatures.SupportsAVX = true;
@@ -532,23 +664,6 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
// Set FPCR back to original just in case anything changed
SetFPCR(OriginalFPCR);
if (HostFeatures.SupportsRAND) {
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
const uint32_t MIDR_Implementer = (MIDR >> 24) & 0xFF;
const uint32_t MIDR_PartNum = (MIDR >> 4) & 0xFFF;
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
// - The RNDRRS register will never read a random number. (Always return failure)
// This is contrary to x86 RNG behaviour where it allows spurious failure with RDSEED, but guarantees eventual success.
// This manifested itself on Linux when an x86 processor failed to guarantee forward progress and boot of services would infinite
// loop. Just disable this extension if this CPU is detected.
HostFeatures.SupportsRAND = false;
}
}
#endif
#ifdef VIXL_SIMULATOR
@@ -560,16 +675,17 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
HostFeatures.SupportsSHA = true;
HostFeatures.SupportsPMULL_128Bit = true;
HostFeatures.SupportsAES256 = true;
// Simulator doesn't support these
HostFeatures.SupportsRPRES = false;
HostFeatures.SupportsAFP = false;
#else
// Check if we can support cacheline clears
uint32_t DCZID = GetDCZID();
if ((DCZID & DCZID_DZP_MASK) == 0) {
uint32_t DCZID_Log2 = DCZID & DCZID_BS_MASK;
uint32_t DCZID_Bytes = (1 << DCZID_Log2) * sizeof(uint32_t);
if (Features.GetDCZID().SupportsDCZVA()) {
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
constexpr static uint64_t CACHELINE_SIZE = 64;
HostFeatures.SupportsCLZERO = DCZID_Bytes == CACHELINE_SIZE;
HostFeatures.SupportsCLZERO = Features.GetDCZID().BlockSizeInBytes() == CACHELINE_SIZE;
}
#endif
@@ -598,21 +714,28 @@ FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool Support
#endif
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
HandleErrata(&HostFeatures, MIDR);
OverrideFeatures(&HostFeatures, ForceSVEWidth());
return HostFeatures;
}
FEXCore::HostFeatures FetchHostFeatures() {
#ifdef _M_X86_64
CPUFeatures Features = CPUFeaturesAll {};
FEX_CONFIG_OPT(CPUFeatureRegisters, CPUFEATUREREGISTERS);
// Vixl simulator doesn't support AFP.
Features.RemoveFeature(CPUFeatures::Feature::AFP);
// Vixl simulator doesn't support RPRES.
Features.RemoveFeature(CPUFeatures::Feature::RPRES);
CPUFeatures Features {};
if (!CPUFeatureRegisters().empty()) {
Features = GetCPUFeaturesFromConfig(CPUFeatureRegisters());
} else {
#ifdef _M_X86_64
Features = CPUFeaturesAll {};
// Vixl simulator doesn't support AFP.
Features.RemoveFeature(CPUFeatures::Feature::AFP);
// Vixl simulator doesn't support RPRES.
Features.RemoveFeature(CPUFeatures::Feature::RPRES);
#else
CPUFeatures Features = GetCPUFeaturesFromIDRegisters();
Features = GetCPUFeaturesFromIDRegisters();
#endif
}
uint64_t CTR = 0;
uint64_t MIDR = 0;
@@ -623,8 +746,9 @@ FEXCore::HostFeatures FetchHostFeatures() {
__asm volatile("mrs %[midr], midr_el1" : [midr] "=r"(MIDR));
#endif
auto HostFeatures = FetchHostFeatures(Features, true, CTR, MIDR);
FEXCore::HostFeatures HostFeatures = {};
FillMIDRInformationViaLinux(&HostFeatures);
FetchHostFeatures(Features, HostFeatures, true, CTR, MIDR);
HostFeatures.SupportsCPUIndexInTPIDRRO = false;
return HostFeatures;
+69 -30
View File
@@ -8,6 +8,23 @@
namespace FEX {
class CPUFeatures {
public:
class FeatureReg {
public:
void SetReg(uint64_t _Reg) {
Reg = _Reg;
}
uint64_t Get() const {
return Reg;
}
protected:
// All feature flag fields are 4-bits.
uint64_t GetField(uint64_t Offset) const {
return (Reg >> Offset) & 0b1111;
}
uint64_t Reg {};
};
enum class Feature : uint32_t {
// ISAR0
AES,
@@ -100,23 +117,25 @@ public:
MAX,
};
static_assert(FEXCore::ToUnderlying(Feature::MAX) < 128);
static_assert((FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) == 1);
class DCZIDReg final : public FeatureReg {
public:
bool SupportsDCZVA() const {
return (Reg & DCZID_DZP_MASK) == 0;
}
bool Supports(Feature feat) const {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
return (FeatureBits[DWordSelect] >> BitSelect) & 1;
}
uint32_t BlockSizeInBytes() const {
uint32_t DCZID_Log2 = Reg & DCZID_BS_MASK;
return (1 << DCZID_Log2) * sizeof(uint32_t);
}
void RemoveFeature(Feature feat) {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
FeatureBits[DWordSelect] &= ~(1ULL << BitSelect);
}
protected:
void FillFeatureFlags();
private:
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
[[maybe_unused]] constexpr static uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
[[maybe_unused]] constexpr static uint32_t DCZID_BS_MASK = 0b0'1111;
};
// This list is informed by Linux kernel's `Documentation/arch/arm64/cpu-feature-registers.rst`
enum class FeatureRegType {
@@ -132,24 +151,11 @@ protected:
ISAR2_EL1,
};
class FeatureReg {
public:
void SetReg(uint64_t _Reg) {
Reg = _Reg;
}
protected:
// All feature flag fields are 4-bits.
uint64_t GetField(uint64_t Offset) const {
return (Reg >> Offset) & 0b1111;
}
uint64_t Reg {};
};
#define FIELD_FETCHER(feature, field, minimum_field) \
bool Supports##feature() const { \
return GetField(field) >= minimum_field; \
}
class ISAR0Reg final : public FeatureReg {
public:
FIELD_FETCHER(AES, AES, 0b0001);
@@ -601,8 +607,11 @@ protected:
ATS1A = 15 * 4,
};
};
class SVEVLReg final : public FeatureReg {};
#undef FIELD_FETCHER
ISAR0Reg ISAR0;
PFR0Reg PFR0;
PFR1Reg PFR1;
@@ -613,6 +622,34 @@ protected:
MMFR2Reg MMFR2;
MMFR1Reg MMFR1;
ISAR2Reg ISAR2;
DCZIDReg DCZID;
SVEVLReg SVEVL;
static_assert(FEXCore::ToUnderlying(Feature::MAX) < 128);
static_assert((FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) == 1);
bool Supports(Feature feat) const {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
return (FeatureBits[DWordSelect] >> BitSelect) & 1;
}
void RemoveFeature(Feature feat) {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
FeatureBits[DWordSelect] &= ~(1ULL << BitSelect);
}
const DCZIDReg& GetDCZID() const {
return DCZID;
}
uint64_t GetSVEVectorLengthInBits() const {
return SVEVL.Get();
}
protected:
void FillFeatureFlags();
uint64_t FeatureBits[(FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) + 1] {};
@@ -625,6 +662,8 @@ protected:
void FillMIDRInformationViaLinux(FEXCore::HostFeatures* Features);
FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool SupportsCacheMaintenanceOps, uint64_t CTR, uint64_t MIDR);
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
uint64_t MIDR);
FEXCore::HostFeatures FetchHostFeatures();
FEX::CPUFeatures GetCPUFeaturesFromIDRegisters();
} // namespace FEX
+33
View File
@@ -0,0 +1,33 @@
add_executable(FEXCompatTool
CompatTool.cpp)
target_link_libraries(FEXCompatTool
PRIVATE
FEXCore Common CommonTools JemallocLibs)
install(TARGETS FEXCompatTool
RUNTIME
DESTINATION /
COMPONENT Runtime
)
add_executable(FEXServerManager
ServerManager.cpp)
target_link_libraries(FEXServerManager
PRIVATE
FEXCore Common CommonTools JemallocLibs)
install(TARGETS FEXServerManager
RUNTIME
DESTINATION bin
COMPONENT Runtime
)
# Description json gets installed into root of depot
install(FILES emulator.json
DESTINATION /
COMPONENT Runtime)
install(FILES ConfigTemplate.json
DESTINATION /
COMPONENT Runtime)
+197
View File
@@ -0,0 +1,197 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: Bin|FEXCompatTool
desc: Used for launching games from Steam
$end_info$
*/
#include "PortabilityInfo.h"
#include "Common/Config.h"
#include "FEXCore/Utils/FileLoading.h"
#include "FEXCore/Utils/StringUtils.h"
#include "FEXHeaderUtils/Filesystem.h"
#include <stdlib.h>
#include <tiny-json.h>
fextl::string GenerateSteamConfigTemplate(const FEX::Config::PortableInformation& PortableInfo) {
const auto ConfigTemplatePath = PortableInfo.InterpreterPath + "ConfigTemplate.json";
if (!FHU::Filesystem::Exists(ConfigTemplatePath)) {
return {};
}
fextl::string Data;
if (!FEXCore::FileLoading::LoadFile(Data, ConfigTemplatePath)) {
return {};
}
// Try and find a mount point.
fextl::string MountPoint {};
const char* RuntimeDir = getenv("XDG_RUNTIME_DIR");
if (RuntimeDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", RuntimeDir);
} else {
const auto UserDirectory = fextl::fmt::format("/run/user/{}", geteuid());
if (FHU::Filesystem::Exists(UserDirectory)) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", UserDirectory);
} else {
const char* CacheDir = getenv("XDG_CACHE_HOME");
if (CacheDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", CacheDir);
} else {
// We tried really hard to find a mount path.
MountPoint = "~/.cache/fexrootfs/";
}
}
}
// Update the @FEX_COMPAT_TOOL@ config to point to the root of the depot.
FEXCore::StringUtils::ReplaceAllInPlace(Data, "@FEX_COMPAT_TOOL@", PortableInfo.InterpreterPath);
// TODO: This path is getting phased out.
FEXCore::StringUtils::ReplaceAllInPlace(Data, "@FEX_ROOTFS_PATH@", MountPoint);
// Save the json.
const auto ConfigPath = FEX::Config::GetConfigDirectory(false, PortableInfo);
const auto ConfigLocation = ConfigPath + "Config.json";
if (!FHU::Filesystem::CreateDirectories(ConfigPath)) {
return {};
}
auto File = FEXCore::File::File(ConfigLocation.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
if (!File.IsValid()) {
return {};
}
File.Write(Data.data(), Data.size());
return ConfigPath;
}
fextl::string GenerateSteamAppConfig(const FEX::Config::PortableInformation& PortableInfo) {
const auto user_config = getenv("FEX_APP_CONFIG");
if (user_config) {
// If user supplied config then don't use Steam config.
return {};
}
// Current supported Steam options.
struct SteamOptions {
bool TSO = true;
bool Multiblock = true;
bool Thunks_GL = false;
bool Thunks_Vulkan = false;
bool EnableLogging = false;
};
SteamOptions Options {};
// Game overrides.
const auto steam_fex_tso = getenv("STEAM_FEX_TSOENABLED");
if (steam_fex_tso) {
Options.TSO = std::strtoull(steam_fex_tso, nullptr, 0) != 0;
}
const auto steam_fex_multiblock = getenv("STEAM_FEX_MULTIBLOCK");
if (steam_fex_multiblock) {
Options.Multiblock = std::strtoull(steam_fex_multiblock, nullptr, 0) != 0;
}
const auto steam_fex_logging = getenv("STEAM_FEX_LOG");
if (steam_fex_logging) {
Options.EnableLogging = std::strtoull(steam_fex_logging, nullptr, 0) != 0;
}
// UI overrides.
const auto steam_fex_compat = getenv("STEAM_COMPAT_FEX_CONFIG");
if (steam_fex_compat) {
const auto steam_fex_compat_view = std::string_view(steam_fex_compat);
if (steam_fex_compat_view.find("TSOEnabled:1") != steam_fex_compat_view.npos) {
Options.TSO = true;
}
if (steam_fex_compat_view.find("Multiblock:1") != steam_fex_compat_view.npos) {
Options.Multiblock = true;
}
if (steam_fex_compat_view.find("ThunksDB_GL:1") != steam_fex_compat_view.npos) {
Options.Thunks_GL = true;
}
if (steam_fex_compat_view.find("ThunksDB_Vulkan:1") != steam_fex_compat_view.npos) {
Options.Thunks_Vulkan = true;
}
}
// Create the json.
char Buffer[4096];
char* Dest {};
Dest = json_objOpen(Buffer, nullptr);
{
Dest = json_objOpen(Dest, "Config");
Dest = json_str(Dest, "TSOEnabled", Options.TSO ? "1" : "0");
Dest = json_str(Dest, "Multiblock", Options.Multiblock ? "1" : "0");
Dest = json_str(Dest, "SilentLog", Options.EnableLogging ? "0" : "1");
if (Options.EnableLogging) {
Dest = json_str(Dest, "OutputLog", "server");
}
Dest = json_objClose(Dest);
}
{
Dest = json_objOpen(Dest, "ThunksDB");
Dest = json_str(Dest, "GL", Options.Thunks_GL ? "1" : "0");
Dest = json_str(Dest, "Vulkan", Options.Thunks_Vulkan ? "1" : "0");
Dest = json_objClose(Dest);
}
Dest = json_objClose(Dest);
json_end(Dest);
// Save the json.
const auto ConfigPath = FEX::Config::GetConfigDirectory(false, PortableInfo);
const auto ConfigLocation = ConfigPath + "app_config.json";
if (!FHU::Filesystem::CreateDirectories(ConfigPath)) {
return {};
}
auto File = FEXCore::File::File(ConfigLocation.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
if (!File.IsValid()) {
return {};
}
File.Write(Buffer, strlen(Buffer));
return ConfigLocation;
}
int main(int argc, const char** argv) {
const auto PortableInfo = FEX::ReadPortabilityInformation();
const auto TemplateConfigPath = GenerateSteamConfigTemplate(PortableInfo);
const auto AppConfigPath = GenerateSteamAppConfig(PortableInfo);
if (!TemplateConfigPath.empty()) {
setenv("FEX_APP_CONFIG_LOCATION", TemplateConfigPath.c_str(), true);
}
if (!AppConfigPath.empty()) {
setenv("FEX_APP_CONFIG", AppConfigPath.c_str(), true);
}
const auto FEXInterpreterPath = PortableInfo.InterpreterPath + "usr/bin/FEX";
// Due to no arguments for this application, just replace argv[0] and execve again.
argv[0] = FEXInterpreterPath.c_str();
execv(FEXInterpreterPath.c_str(), const_cast<char* const*>(argv));
// Save errno as it can change after calling `perror`.
const auto saved_errno = errno;
perror(argv[0]);
if (saved_errno == ENOENT) {
return 127;
}
return 126;
}
+9
View File
@@ -0,0 +1,9 @@
{
"Config": {
"X87ReducedPrecision": "1",
"RootFS": "@FEX_ROOTFS_PATH@/",
"ThunkHostLibs": "@FEX_COMPAT_TOOL@/usr/lib/aarch64-linux-gnu/fex-emu/HostThunks",
"ThunkGuestLibs": "@FEX_COMPAT_TOOL@/usr/share/fex-emu/GuestThunks",
"ProfileStats": "1"
}
}
+96
View File
@@ -0,0 +1,96 @@
// SPDX-License-Identifier: MIT
#include "PortabilityInfo.h"
#include "Common/FEXServerClient.h"
#include <cstdio>
#include <unistd.h>
#include <poll.h>
void MsgHandler(LogMan::DebugLevels Level, const char* Message) {
const auto Style = fmt::text_style {};
const auto Output = fextl::fmt::format("{} {}\n", fmt::styled(LogMan::DebugLevelStr(Level), Style), Message);
write(STDERR_FILENO, Output.c_str(), Output.size());
fsync(STDERR_FILENO);
}
void AssertHandler(const char* Message) {
return MsgHandler(LogMan::ASSERT, Message);
}
void SignalPVToContinue() {
// Tell pressure-vessel that the startup was a success.
const auto ReadyMsg = "READY=1\n";
write(STDOUT_FILENO, ReadyMsg, strlen(ReadyMsg));
// pressure-vessel is waiting for EOF on STDOUT from this process to ensure it can run FEX processes.
// dup2 atomically replaces stdout with a copy of stderr to achieve this.
dup2(STDERR_FILENO, STDOUT_FILENO);
}
struct PipesType {
int read_pipe {-1};
int write_pipe {-1};
};
PipesType get_pipe() {
PipesType pipes {};
pipe(&pipes.read_pipe);
return pipes;
}
int main(int argc, const char** argv, char** const envp) {
LogMan::Throw::InstallHandler(AssertHandler);
LogMan::Msg::InstallHandler(MsgHandler);
const auto PortableInfo = FEX::ReadPortabilityInformation();
FEX::Config::LoadConfig({}, envp, PortableInfo);
// Reload the meta layer
FEXCore::Config::ReloadMetaLayer();
auto pipes = get_pipe();
// Set the write side to close on exec.
fcntl(pipes.write_pipe, F_SETFD, FD_CLOEXEC);
// Give the read end of the pipe to FEXServer.
auto ServerFD = FEXServerClient::StartServer(PortableInfo.InterpreterPath, pipes.read_pipe);
if (ServerFD == -1) {
perror("Couldn't start FEXServer");
return 126;
}
// FEXServer is now running. Tell PV to continue.
SignalPVToContinue();
// Don't need the read pipe anymore.
close(pipes.read_pipe);
pipes.read_pipe = -1;
// Now that the server is started and watching our pipe, we can close the returned FD, as it'll stay open as long as the pipe is open.
close(ServerFD);
ServerFD = -1;
// stdin will be a pipe, so wait until that FD is closed.
while (true) {
pollfd p {
.fd = STDIN_FILENO,
.events = POLLRDHUP,
.revents = 0,
};
int events = poll(&p, 1, -1);
if (events == -1 && errno == EINTR) {
continue;
}
if (events > 0 && (p.revents & (POLLRDHUP | POLLERR | POLLHUP | POLLNVAL))) {
// Error or pressure-vessel hung-up.
break;
}
}
// Terminating will clean-up.
return 0;
}
+13
View File
@@ -0,0 +1,13 @@
{
"emulator_v0": {
"argv": "./usr/bin/FEX",
"environment": { "FEX_PORTABLE": "1" },
"container_argv": "./usr/bin/FEX",
"container_environment": { "FEX_ROOTFS": "" },
"main_argv": "./FEXCompatTool",
"server_argv": "./usr/bin/FEXServerManager",
"emulated_architectures": ["x86_64-linux-gnu", "i386-linux-gnu"],
"required_architectures": ["aarch64-linux-gnu"],
"required_libraries": ["libc.so.6", "libstdc++.so.6"]
}
}
+5 -1
View File
@@ -14,10 +14,14 @@ if (NOT MINGW_BUILD)
add_subdirectory(FEXGDBReader/)
endif()
add_subdirectory(FEXRootFSFetcher/)
if (NOT BUILD_STEAM_SUPPORT)
add_subdirectory(FEXRootFSFetcher/)
endif()
add_subdirectory(FEXGetConfig/)
add_subdirectory(FEXServer/)
add_subdirectory(FEXBash/)
add_subdirectory(FEXOfflineCompiler/)
add_subdirectory(CodeSizeValidation/)
add_subdirectory(LinuxEmulation/)
+2 -2
View File
@@ -52,8 +52,8 @@ public:
{
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSection(CTX->GetCodeInvalidationMutex(), Thread);
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
CTX->InvalidateGuestCodeRange(Thread, Accumulator, reinterpret_cast<uint64_t>(CodeStart), MAX_CODE_SIZE);
CTX->InvalidateCodeBuffersCodeRange(reinterpret_cast<uint64_t>(CodeStart), MAX_CODE_SIZE);
CTX->InvalidateThreadCachedCodeRange(Thread, reinterpret_cast<uint64_t>(CodeStart), MAX_CODE_SIZE);
}
ClearStats();
+1
View File
@@ -5,6 +5,7 @@
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <unistd.h>
namespace FEX {
+5 -1
View File
@@ -78,6 +78,10 @@ void ConfigModel::Reload() {
if (!LoadedConfig->OptionExists(Option.first)) {
continue;
}
if (std::holds_alternative<fextl::list<fextl::string>>(Option.second)) {
// Omit string lists from the model since they require special handling
continue;
}
auto& [Name, TypeId] = ConfigToNameLookup.find(Option.first)->second;
auto Item = new QStandardItem(QString::fromStdString(Name));
@@ -122,7 +126,7 @@ void ConfigModel::setStringList(const QString& Name, const QStringList& Values)
const auto& Option = NameToConfigLookup.at(Name.toStdString());
LoadedConfig->Erase(Option);
for (auto& Value : Values) {
LoadedConfig->Set(Option, Value.toStdString().c_str());
LoadedConfig->AppendStrArrayValue(Option, Value.toStdString().c_str());
}
Reload();
}
+25
View File
@@ -2,6 +2,7 @@
#include "Common/cpp-optparse/OptionParser.h"
#include "Common/Config.h"
#include "Common/FEXServerClient.h"
#include "Common/HostFeatures.h"
#include "git_version.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/PrctlUtils.h>
@@ -112,6 +113,10 @@ int main(int argc, char** argv, char** envp) {
Parser.add_option("--tso-emulation-info").action("store_true").help("Print how FEX is emulating the x86-TSO memory model.");
#ifdef _M_ARM_64
Parser.add_option("--identification-reg-info").action("store_true").help("Print identification registers");
#endif
Parser.add_option("--version").action("store_true").help("Print the installed FEX-Emu version");
optparse::Values Options = Parser.parse_args(argc, argv);
@@ -230,5 +235,25 @@ int main(int argc, char** argv, char** envp) {
fprintf(stdout, "\t64-Byte strict split-lock emulation: %s\n", StrictInProcessSplitLocks() ? "In-process mutex" : "Tearing");
}
#ifdef _M_ARM_64
if (Options.is_set_by_user("identification_reg_info")) {
auto Features = FEX::GetCPUFeaturesFromIDRegisters();
fextl::string features {};
features += fmt::format("isar0=0x{:x},", Features.ISAR0.Get());
features += fmt::format("isar1=0x{:x},", Features.ISAR1.Get());
features += fmt::format("isar2=0x{:x},", Features.ISAR2.Get());
features += fmt::format("pfr0=0x{:x},", Features.PFR0.Get());
features += fmt::format("pfr1=0x{:x},", Features.PFR1.Get());
features += fmt::format("midr=0x{:x},", Features.MIDR.Get());
features += fmt::format("mmfr0=0x{:x},", Features.MMFR0.Get());
features += fmt::format("mmfr1=0x{:x},", Features.MMFR1.Get());
features += fmt::format("mmfr2=0x{:x},", Features.MMFR2.Get());
features += fmt::format("zfr0=0x{:x},", Features.ZFR0.Get());
features += fmt::format("dczid=0x{:x},", Features.DCZID.Get());
features += fmt::format("svevl=0x{:x}", Features.SVEVL.Get());
fprintf(stderr, "Features: '%s'\n", features.c_str());
}
#endif
return 0;
}
+77 -6
View File
@@ -17,6 +17,7 @@
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/fextl/list.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
@@ -29,7 +30,9 @@
#include <sys/auxv.h>
#include <sys/mman.h>
#include <sys/personality.h>
#include <sys/prctl.h>
#include <sys/random.h>
#include <linux/prctl.h>
#define PAGE_START(x) ((x) & ~(uintptr_t)(4095))
#define PAGE_OFFSET(x) ((x) & 4095)
@@ -73,7 +76,7 @@ class ELFCodeLoader final : public FEX::CodeLoader {
return FEXCore::AlignUp(max_map_address - min_map_address, FEXCore::Utils::FEX_PAGE_SIZE);
}
bool MapFile(const ELFParser& file, uintptr_t Base, const Elf64_Phdr& Header, int prot, int flags, FEX::HLE::SyscallHandler* const Handler) {
bool MapFile(const ELFParser& file, uintptr_t Base, const Elf64_Phdr& Header, int prot, int flags, FEX::HLE::SyscallMmapInterface* const Handler) {
auto addr = Base + PAGE_START(Header.p_vaddr);
auto size = Header.p_filesz + PAGE_OFFSET(Header.p_vaddr);
@@ -121,7 +124,7 @@ class ELFCodeLoader final : public FEX::CodeLoader {
return rv;
}
std::optional<uintptr_t> LoadElfFile(ELFParser& Elf, uintptr_t* BrkBase, FEX::HLE::SyscallHandler* const Handler, uint64_t LoadHint = 0) {
std::optional<uintptr_t> LoadElfFile(ELFParser& Elf, uintptr_t* BrkBase, FEX::HLE::SyscallMmapInterface* const Handler, uint64_t LoadHint = 0) {
uintptr_t LoadBase = 0;
uintptr_t BrkLoadBase = 0;
@@ -254,7 +257,7 @@ public:
ELFCodeLoader(const fextl::string& Filename, int ProgramFDFromEnv, const fextl::string& RootFS, const fextl::vector<fextl::string>& args,
const fextl::vector<fextl::string>& ParsedArgs, char** const envp = nullptr,
FEXCore::Config::Value<FEXCore::Config::DefaultValues::Type::StringArrayType>* AdditionalEnvp = nullptr) {
FEXCore::Config::Value<FEXCore::Config::StringArrayType>* AdditionalEnvp = nullptr, bool SkipInterpreter = false) {
ApplicationArgs = args;
bool LoadedWithFD = false;
@@ -302,7 +305,7 @@ public:
const auto AdditionalArgs = AdditionalArguments.All();
ApplicationArgs.insert(ApplicationArgs.end(), AdditionalArgs.begin(), AdditionalArgs.end());
if (!MainElf.InterpreterElf.empty()) {
if (!MainElf.InterpreterElf.empty() && !SkipInterpreter) {
if (!InterpElf.ReadElf(ResolveRootfsFile(MainElf.InterpreterElf, RootFS)) && !InterpElf.ReadElf(MainElf.InterpreterElf)) {
return;
}
@@ -374,7 +377,11 @@ public:
uint64_t val;
};
bool MapMemory(FEX::HLE::SyscallHandler* const Handler) {
std::optional<uintptr_t> LoadMainElfFile(uintptr_t* BrkBase, FEX::HLE::SyscallMmapInterface* const Handler, uint64_t LoadHint = 0) {
return LoadElfFile(MainElf, BrkBase, Handler, LoadHint);
}
bool MapMemory(FEX::HLE::SyscallMmapInterface* const Handler) {
for (const auto& Header : MainElf.phdrs) {
if (Header.p_type == PT_GNU_STACK) {
if (Header.p_flags & PF_X) {
@@ -707,6 +714,63 @@ public:
*AuxTabSize = sizeof(AuxType) * AuxVariables.size();
}
// Get the current memory map from /proc/self/stat
static bool GetCurrentMap(struct prctl_mm_map& map) {
// /proc/self/stat has 52 fields of at most 20 digits each (UINT64_MAX).
// 52*20 = 1040, so 2048 is a conservative upper bound
char stat_buffer[2048];
ssize_t bytes_read = FEXCore::FileLoading::LoadFileToBuffer("/proc/self/stat", stat_buffer);
// Ensure we don't read past the end into garbage data
stat_buffer[std::clamp(bytes_read, 0L, static_cast<ssize_t>(sizeof(stat_buffer)) - 1)] = '\0';
// See man proc_pid_stat
int items_read = sscanf(stat_buffer,
"%*d %*s %*c %*d %*d " // 1 to 5
"%*d %*d %*d %*u %*u " // 6 to 10
"%*u %*u %*u %*u %*u " // 11 to 15
"%*d %*d %*d %*d %*d " // 16 to 20
"%*d %*u %*u %*d %*u " // 21 to 25
"%llu %llu %llu %*u %*u " // 26 to 30
"%*u %*u %*u %*u %*u " // 31 to 35
"%*u %*u %*d %*d %*u " // 36 to 40
"%*u %*u %*u %*d %llu " // 40 to 45
"%llu %llu %llu %llu %llu " // 46 to 50
"%llu", // 51
&map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
&map.arg_start, &map.arg_end, &map.env_start, &map.env_end);
if (items_read != 10) {
return false;
}
map.brk = reinterpret_cast<uint64_t>(sbrk(0));
// The kernel will leave these values unchanged, see implementation in sys.c
map.auxv = NULL;
map.auxv_size = 0;
map.exe_fd = -1;
return true;
}
// Point the OS to our new stack's argument data
void RemapArgumentData(uintptr_t NewArgStart, uint64_t ArgSize) {
struct prctl_mm_map map {};
if (GetCurrentMap(map)) {
map.arg_start = NewArgStart;
map.arg_end = NewArgStart + ArgSize;
int r = prctl(PR_SET_MM, PR_SET_MM_MAP, &map, sizeof(map), 0L);
if (r != 0) {
LogMan::Msg::EFmt("Failed to remap /proc/pid/cmdline data (prctl failed: result {}, errno {})", r, errno);
}
} else {
LogMan::Msg::EFmt("Failed to remap /proc/pid/cmdline data (GetCurrentMap failed)");
}
}
// Setups the stack initial data (argv, envp, auxv)
void SetupStack() {
StackPointer += StackSize();
@@ -728,7 +792,7 @@ public:
TotalArgumentMemSize += sizeof(auxv32_t) * AuxVariables.size();
}
uint64_t ArgumentOffset = TotalArgumentMemSize;
ArgumentOffset = TotalArgumentMemSize;
TotalArgumentMemSize += ArgumentBackingSize;
uint64_t EnvpOffset = TotalArgumentMemSize;
@@ -803,6 +867,8 @@ public:
SetupPointers<uint32_t, auxv32_t, 4>(StackPointer, AuxVOffset, ArgumentOffset, EnvpOffset, ApplicationArgs, EnvironmentVariables,
AuxVariables, &AuxTabBase, &AuxTabSize);
}
RemapArgumentData(StackPointer + ArgumentOffset, ArgumentBackingSize);
}
fextl::vector<const char*> GetExecveArguments() const override {
@@ -820,6 +886,10 @@ public:
return BaseOffset;
}
uint64_t GetMainElfBase() const {
return MainElfBase;
}
bool Is64BitMode() const {
return MainElf.type == ::ELFLoader::ELFContainer::TYPE_X86_64;
}
@@ -910,6 +980,7 @@ public:
fextl::list<auxv_t> AuxVariables;
uint64_t AuxTabBase {}, AuxTabSize {};
uint64_t ArgumentBackingSize {};
uint64_t ArgumentOffset {};
uint64_t EnvironmentBackingSize {};
uint64_t BaseOffset {};
void* VDSOBase {};
+33 -5
View File
@@ -222,6 +222,26 @@ void CheckForGCS() {
}
} // namespace FEX::GCS
namespace FEX::UnalignedAtomic {
void SetupKernelUnalignedAtomics() {
#ifndef PR_ARM64_SET_UNALIGN_ATOMIC
#define PR_ARM64_SET_UNALIGN_ATOMIC 0x46455849
#define PR_ARM64_UNALIGN_ATOMIC_EMULATE (1UL << 0)
#define PR_ARM64_UNALIGN_ATOMIC_BACKPATCH (1UL << 1)
#define PR_ARM64_UNALIGN_ATOMIC_STRICT_SPLIT_LOCKS (1UL << 2)
#endif
// Interfaces with downstream FEX kernel patches to control unaligned atomic handling
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
FEX_CONFIG_OPT(KernelUnalignedAtomicBackpatching, KERNELUNALIGNEDATOMICBACKPATCHING);
uint64_t Flags = (StrictInProcessSplitLocks() ? PR_ARM64_UNALIGN_ATOMIC_STRICT_SPLIT_LOCKS : 0) |
(KernelUnalignedAtomicBackpatching() ? PR_ARM64_UNALIGN_ATOMIC_BACKPATCH : 0) | PR_ARM64_UNALIGN_ATOMIC_EMULATE;
prctl(PR_ARM64_SET_UNALIGN_ATOMIC, Flags, 0, 0, 0);
}
} // namespace FEX::UnalignedAtomic
/**
* @brief Get an FD from an environment variable and then unset the environment variable.
*
@@ -331,7 +351,7 @@ int main(int argc, char** argv, char** const envp) {
if (LogFD == -1) {
LogMan::Msg::EFmt("Couldn't open log file. Going Silent.");
SilentLog = true;
::SilentLog = true;
} else {
OutputFD = LogFD;
}
@@ -408,6 +428,10 @@ int main(int argc, char** argv, char** const envp) {
char* RealPath = realpath(Program.ProgramPath.c_str(), ExistsTempPath);
if (RealPath) {
FEXCore::Config::Set(FEXCore::Config::CONFIG_APP_FILENAME, fextl::string(RealPath));
} else {
// Can happen when jumping in to pressure-vessel.
// `/usr/lib/pressure-vessel/from-host/libexec/steam-runtime-tools-0/pv-adverb` can't get resolved.
FEXCore::Config::Set(FEXCore::Config::CONFIG_APP_FILENAME, Program.ProgramPath);
}
}
FEXCore::Config::Set(FEXCore::Config::CONFIG_APP_CONFIG_NAME, Program.ProgramName);
@@ -458,6 +482,7 @@ int main(int argc, char** argv, char** const envp) {
// Setup TSO hardware emulation immediately after initializing the context.
FEX::TSO::SetupTSOEmulation(CTX.get());
FEX::UnalignedAtomic::SetupKernelUnalignedAtomics();
if (!Loader.Is64BitMode()) {
// Tell the kernel we want to use the compat input syscalls even though we're
@@ -475,10 +500,17 @@ int main(int argc, char** argv, char** const envp) {
auto SyscallHandler = Loader.Is64BitMode() ?
FEX::HLE::x64::CreateHandler(CTX.get(), SignalDelegation.get(), ThunkHandler.get()) :
FEX::HLE::x32::CreateHandler(CTX.get(), SignalDelegation.get(), ThunkHandler.get(), std::move(Allocator));
if (FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
CTX->SetCodeMapWriter(fextl::make_unique<FEXCore::CodeMapWriter>(*SyscallHandler));
}
// Load VDSO in to memory prior to mapping our ELFs.
auto VDSOMapping = FEX::VDSO::LoadVDSOThunks(Loader.Is64BitMode(), SyscallHandler.get());
// Pass in our VDSO thunks
ThunkHandler->AppendThunkDefinitions(FEX::VDSO::GetVDSOThunkDefinitions(Loader.Is64BitMode()));
SignalDelegation->SetVDSOSymbols();
// Now that we have the syscall handler. Track some FDs that are FEX owned.
if (OutputFD > 2) {
SyscallHandler->FM.TrackFEXFD(OutputFD);
@@ -524,10 +556,6 @@ int main(int argc, char** argv, char** const envp) {
SignalDelegation->RegisterTLSState(ParentThread);
ThunkHandler->RegisterTLSState(ParentThread);
// Pass in our VDSO thunks
ThunkHandler->AppendThunkDefinitions(FEX::VDSO::GetVDSOThunkDefinitions(Loader.Is64BitMode()));
SignalDelegation->SetVDSOSigReturn();
SyscallHandler->DeserializeSeccompFD(ParentThread, FEXSeccompFD);
CTX->ExecuteThread(ParentThread->Thread);
@@ -0,0 +1,25 @@
add_executable(FEXOfflineCompiler Main.cpp)
target_link_libraries(FEXOfflineCompiler
PRIVATE
Common
CommonTools
cpp-optparse
FEXCore
JemallocLibs
LinuxEmulation
${PTHREAD_LIB}
fmt::fmt)
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(FEXOfflineCompiler
PRIVATE
"LINKER:--gc-sections"
"LINKER:--strip-all"
"LINKER:--as-needed")
endif()
install(TARGETS FEXOfflineCompiler
RUNTIME
DESTINATION bin
COMPONENT Runtime)
+280
View File
@@ -0,0 +1,280 @@
// SPDX-License-Identifier: MIT
#include "../FEXInterpreter/ELFCodeLoader.h"
#include <DummyHandlers.h>
#include <PortabilityInfo.h>
#include <Thunks.h>
#include <FEXCore/Core/CodeCache.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/HostFeatures.h>
#include <Common/ArgumentLoader.h>
#include <Common/Config.h>
#include <Common/FEXServerClient.h>
#include <Common/HostFeatures.h>
#include <OptionParser.h>
#include <fmt/printf.h>
#include <fstream>
#include <optional>
class AOTSyscallHandler : public FEXCore::HLE::SyscallHandler, public FEX::HLE::SyscallMmapInterface {
public:
AOTSyscallHandler(FEXCore::HLE::SyscallOSABI SyscallOSABI) {
OSABI = SyscallOSABI;
}
uint64_t HandleSyscall(FEXCore::Core::CpuStateFrame* Frame, FEXCore::HLE::SyscallArguments* Args) override {
// Don't do anything
return 0;
}
FEXCore::ExecutableFileInfo FileInfo;
uintptr_t VAFileStart = 0;
// These are no-ops implementations of the SyscallHandler API
std::optional<FEXCore::ExecutableFileSectionInfo> LookupExecutableFileSection(FEXCore::Core::InternalThreadState&, uint64_t) override {
return FEXCore::ExecutableFileSectionInfo {FileInfo, VAFileStart};
}
FEXCore::HLE::ExecutableRangeInfo QueryGuestExecutableRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Address) override {
return {0, UINT64_MAX, true};
}
void* GuestMmap(FEXCore::Core::InternalThreadState*, void* addr, size_t Size, int prot, int Flags, int fd, off_t offset) override {
auto Ret = mmap(addr, Size, prot, Flags, fd, offset);
if (Ret != MAP_FAILED && VAFileStart == 0) {
VAFileStart = reinterpret_cast<uintptr_t>(Ret);
}
return Ret;
}
uint64_t GuestMunmap(FEXCore::Core::InternalThreadState*, void* addr, uint64_t length) override {
return munmap(addr, length);
}
};
static void MsgHandler(LogMan::DebugLevels Level, const char* Message) {
fmt::print("[{}] {}\n", LogMan::DebugLevelStr(Level), Message);
}
static void AssertHandler(const char* Message) {
fmt::print("[A] {}\n", Message);
}
namespace FEXCore {
inline bool operator<(const ExecutableFileInfo& a, const ExecutableFileInfo& b) noexcept {
return a.FileId < b.FileId;
}
} // namespace FEXCore
template<>
struct std::hash<FEXCore::ExecutableFileInfo> {
std::size_t operator()(const FEXCore::ExecutableFileInfo& Val) const noexcept {
return Val.FileId;
}
};
// Placeholder data to ensure the compile thread doesn't de-reference nullptr data
static FEXCore::Core::CPUState::gdt_segment gdt[32] {};
static FEXCore::Core::InternalThreadState* SetupCompileThread(FEXCore::Context::Context& CTX, bool Is64Bit) {
auto Thread = CTX.CreateThread(0, 0);
auto Frame = Thread->CurrentFrame;
Frame->State.segment_arrays[FEXCore::Core::CPUState::SEGMENT_ARRAY_INDEX_GDT] = &gdt[0];
Frame->State.segment_arrays[FEXCore::Core::CPUState::SEGMENT_ARRAY_INDEX_LDT] = &gdt[0];
Frame->State.cs_idx = FEXCore::Core::CPUState::DEFAULT_USER_CS << 3;
auto GDT = FEXCore::Core::CPUState::GetSegmentFromIndex(Frame->State, Frame->State.cs_idx);
FEXCore::Core::CPUState::SetGDTBase(GDT, 0);
FEXCore::Core::CPUState::SetGDTLimit(GDT, 0xFFFFFU);
Frame->State.cs_cached =
FEXCore::Core::CPUState::CalculateGDTBase(*FEXCore::Core::CPUState::GetSegmentFromIndex(Frame->State, Frame->State.cs_idx));
if (Is64Bit) {
GDT->L = 1; // L = Long Mode = 64-bit
GDT->D = 0; // D = Default Operand SIze = Reserved
} else {
GDT->L = 0; // L = Long Mode = 32-bit
GDT->D = 1; // D = Default Operand Size = 32-bit
}
return Thread;
}
// Returns filename of generated cache on success
static std::optional<std::string>
GenerateSingleCache(const FEXCore::ExecutableFileInfo& Binary, fextl::set<uintptr_t> BlockList, std::string_view OutDir) {
uint64_t CodeCacheConfigId = 0; // TODO: Make unique to active configuration
ELFCodeLoader Loader(Binary.Filename.c_str(), -1, "", fextl::vector<fextl::string> {Binary.Filename.c_str()},
fextl::vector<fextl::string> {}, nullptr, nullptr, true /* skip interpreter */);
if (!Loader.ELFWasLoaded()) {
fmt::print("Invalid or unsupported ELF file.\n");
return std::nullopt;
}
const bool Is64Bit = Loader.Is64BitMode();
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS64BIT_MODE, Is64Bit ? "1" : "0");
// Load HostFeatures
auto HostFeatures = FEX::FetchHostFeatures();
if (!std::filesystem::exists(Binary.Filename)) {
fmt::print("File {} does not exist\n", Binary.Filename);
// TODO: Pressure vessel hits this
return /*EXIT_FAILURE*/ std::nullopt;
}
auto CTX = FEXCore::Context::Context::CreateNewContext(HostFeatures);
auto SignalDelegation = std::make_unique<FEX::DummyHandlers::DummySignalDelegator>();
auto SyscallOSABI = Is64Bit ? FEXCore::HLE::SyscallOSABI::OS_LINUX64 : FEXCore::HLE::SyscallOSABI::OS_LINUX32;
auto SyscallHandler = std::make_unique<AOTSyscallHandler>(SyscallOSABI);
Loader.CalculateHWCaps(CTX.get());
CTX->SetSignalDelegator(SignalDelegation.get());
CTX->SetSyscallHandler(SyscallHandler.get());
auto ThunkHandler = FEX::HLE::CreateThunkHandler();
CTX->SetThunkHandler(ThunkHandler.get());
if (!CTX->InitCore()) {
return std::nullopt;
}
{
if (!Is64Bit) {
// Block upper address space
FEXCore::Allocator::SetupHooks();
}
auto ElfBase = Loader.LoadMainElfFile(nullptr, SyscallHandler.get());
if (!ElfBase.has_value()) {
ERROR_AND_DIE_FMT("Failed to load ELF file {} ({})", Binary.Filename, Binary.FileId);
}
}
auto Thread = SetupCompileThread(*CTX, Is64Bit);
CTX->GetCodeCache().InitiateCacheGeneration();
{
std::vector<std::unique_ptr<ELFCodeLoader>> LoaderMem;
fmt::print(stderr, "Compiling code...\n");
for (auto Addr : BlockList) {
CTX->CompileRIP(Thread, Addr + SyscallHandler->VAFileStart);
}
auto Filename = fmt::format("{}{}-{:016x}", OutDir, FEXCore::CodeMap::GetBaseFilename(Binary, false), CodeCacheConfigId);
auto FilenameNew = Filename + ".new";
int fd = open(FilenameNew.c_str(), O_CREAT | O_WRONLY, 0644);
{
auto Entry = SyscallHandler->LookupExecutableFileSection(*Thread, Loader.GetMainElfBase()).value();
CTX->GetCodeCache().SaveData(*Thread, fd, Entry, 0 /* TODO: Use static base address information if available */);
}
std::filesystem::rename(FilenameNew.c_str(), Filename.c_str());
close(fd);
return Filename;
}
}
// Command handler that parses the given code map and generates a code cache for the selected x86 binary.
// If no binary is selected explicitly, it is inferred from the code map ExecutableFileId block.
static int GenerateCache(int argc, const char** argv) {
optparse::OptionParser Parser {};
Parser.add_option("--outdir").set_default(FEX::Config::GetCacheDirectory() + "cache").help("Output directory for generated cache files");
Parser.add_option("--fileid").help("Select binary to generate cache for");
optparse::Values Options = Parser.parse_args(argc, argv);
if (Parser.args().size() != 1) {
Parser.print_usage();
return 1;
}
const fextl::string CodeMapPath = Parser.args()[0];
std::ifstream Codemap(CodeMapPath.c_str(), std::ios_base::binary);
if (!Codemap) {
fmt::print("Could not open {}\n", CodeMapPath);
return 1;
}
FEXCore::ExecutableFileInfo ProgramName;
std::map<FEXCore::ExecutableFileInfo, fextl::set<uintptr_t>> Data;
{
auto Parsed = FEXCore::CodeMap::ParseCodeMap(Codemap);
// If an explicit file id is selected, use it.
// Otherwise, fall back to an IsExecutable marker (or pick the first entry if there's only one)
auto ExplicitFileId = strtoull(((fextl::string)Options.get("fileid")).data(), nullptr, 16);
if (ExplicitFileId) {
ProgramName.FileId = ExplicitFileId;
ProgramName.Filename = Parsed.at(ExplicitFileId).Filename;
}
for (auto& [FileId, Contents] : Parsed) {
if (!ExplicitFileId && (Contents.IsExecutable || Parsed.size() == 1)) {
ProgramName.FileId = FileId;
ProgramName.Filename = Contents.Filename;
}
Data.emplace(std::piecewise_construct, std::forward_as_tuple(nullptr, FileId, std::move(Contents.Filename)),
std::forward_as_tuple(std::move(Contents.Blocks)));
}
}
if (!ProgramName.FileId) {
fmt::print("Cannot generate cache from unsanitized code map {}", CodeMapPath);
return 1;
}
for (auto& [File, Blocks] : Data) {
if (!Blocks.empty()) {
fmt::print("Parsed {} codemap entries for {} ({:016x})\n", Blocks.size(), File.Filename, File.FileId);
} else {
fmt::print("Found dependency {} ({:016x})\n", File.Filename, File.FileId);
}
}
if (!Data.contains(ProgramName)) {
throw std::runtime_error(fmt::format("Input code map {} did not contain {} ({:016x})", CodeMapPath, ProgramName.Filename, ProgramName.FileId));
}
fextl::string OutDir(Options.get("outdir"));
if (!OutDir.ends_with('/')) {
OutDir.push_back('/');
}
std::filesystem::create_directories(OutDir);
const auto PortableInfo = FEX::ReadPortabilityInformation();
char* envp[] = {nullptr};
FEX::Config::LoadConfig("", envp, PortableInfo);
auto NumBlocks = Data.at(ProgramName).size();
auto GeneratedCache = GenerateSingleCache(ProgramName, Data.at(ProgramName), OutDir);
if (GeneratedCache) {
fmt::print("Successfully populated cache {} ({} blocks) via {}\n\n", GeneratedCache.value(), NumBlocks,
std::filesystem::path {CodeMapPath}.filename().string());
}
return GeneratedCache ? 0 : 1;
}
int main(int argc, char** argv) {
LogMan::Throw::InstallHandler(AssertHandler);
LogMan::Msg::InstallHandler(MsgHandler);
std::vector<const char*> Args {argv + 1, argv + argc};
auto CommandName = std::string {basename(argv[0])} + " " + (argc > 1 ? argv[1] : "");
Args[0] = CommandName.c_str();
if (argc >= 2 && argv[1] == std::string_view {"generate"}) {
return GenerateCache(argc - 1, Args.data());
} else {
fmt::print("Usage: {} <command>\n\n", basename(argv[0]));
fmt::print("Commands:\n");
fmt::print(" generate\tTrigger cache generation from combined code map\n");
return EXIT_FAILURE;
}
}
+3 -1
View File
@@ -247,7 +247,9 @@ void CheckSquashfuse() {
void CheckUnsquashfs() {
const std::array<const char*, 3> ExecveArgs = {
"unsquashfs",
"--help",
// since unsquashfs 4.7.1, -help-all is needed to list decompressors.
// also works with older versions.
"-help-all",
nullptr,
};
Loaded 100 of 186 files, more files were not shown because too many files have changed in this diff. Show more