The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.
With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
ELF headers were read unconditionally because doing so was assumed to be cheap
(as the guest app would read them anyway shortly after). However, relocation
parsing was added since then, which has less predictable performance due to
crossing page boundaries and reading larger amounts of memory. It might be
possible to make the underlying code more efficient, but until that's done
it's better to skip this logic unless needed.
Closes#5390
This allows the context and parent thread objects to be created earlier,
allowing the VDSO and ELFCodeLoader mapping functions to have a thread
object for tracking memory mappings through the regular guest routines.
This means that we no longer need to do any form of deferred handling
for code caching as all the state is ready early in the initialization
process.
A little bit of care needed to be taken to ensure we still close the
ELFCodeLoader's FDs later and that VDSO unmapping happens before tearing
down the parent thread, but overall this is mostly just passing the
InternalThreadState object around as normal.
I couldn't find any functional regression from this change alongside
code caching, but it would be good for @neobrain to double check this.
The ELF tracking thing has an expectation that only portions of ELF
files that are described in the program headers will be mapped
executable. This doesn't hold true as programs will remap random
portions of ELF files as executable. In the case that this occurs, don't
assert out and instead print a warning.
This was discovered as Node.js remaps a portion of itself executable
that isn't described as such in the program headers. I also have a local
unittest that exposes the same problem. I had discovered same problem in
some other program with #5038.
Also fixes a bug where sometimes completely anonymously mapped
executable sneak in and cause a crash, which is kind of silly.
At runtime, glxtest is mapped as follows:
0x000055fd9a030000 0x000055fd9a034000 0x4000 0x0 r--p glxtest
0x000055fd9a034000 0x000055fd9a038000 0x4000 0x3000 r-xp glxtest
0x000055fd9a038000 0x000055fd9a039000 0x1000 0x6000 rw-p glxtest
0x000055fd9a039000 0x000055fd9a03a000 0x1000 0x6000 rw-p glxtest
The problem here is that the last two sections can't be distinguished solely
by their mmap parameters. This would cause the wrong base address to be
inferred for the last mapping. To fix this, we can be more permissive by
allowing multiple candidates to be returned.
In practice, this only affects non-code sections, so it's not a big issue
either way.
The file offset of a file mapping doesn't necessarily match its address
offset in virtual memory from the base file mapping. Indeed, most libraries
violate this assumption.
Now that the MappedResource::FirstVMA reliably identifies the base memory
mapping for a given library (even when that library is mapped multiple times),
this can easily be fixed.
PE/ELF binaries are sometimes mapped multiple times in the same process.
If this happens, there is no longer a unique base virtual address per file.
This breaks assumptions required for code caching: Any time a file mapping
is created, FEX must be able to unambiguously determine the base virtual
address of the mapped library.
This becomes possible by creating a separate MappedResource each time an
ELF header is re-mapped.
Allows applications running entirely under Linux to use the same
extended volatile metadata as Windows.
For example `FEX_EXTENDEDVOLATILEMETADATA=iw4sp.exe\;0xe9da0-0xe9ec7`
this configuration works for both wow64 and Linux to disable the TSO
emulation on the memcpy routine in that game that consumes around 80% of
CPU time in TSO emulation.
Works with Linux native games as well of course.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.