musl build fails due to access to internal glibc member:
```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/SignalDelegator.cpp:655:39: error: no member named '__val' in '__sigset_t'
655 | .SigMask = _context->uc_sigmask.__val[0],
| ~~~~~~~~~~~~~~~~~~~~ ^
1 error generated.
```
add check to determine private glibc or musl member name or throw an error.
fixes: #5461
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
building on clang/musl fails with:
```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/GdbServer.cpp:1174:5: error: use of undeclared identifier 'tgkill'
1174 | tgkill(::getpid(), ::getpid(), SIGKILL);
| ^~~~~~
1 error generated.
```
include and use `FHU::Syscalls::tgkill`.
fixes: #5463
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.
With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
This was causing an unfortunate set of circumstances where Dark Souls
III was modifying MXCSR and we weren't saving it, cause the value to change
from 0x9fc0 to 0.
This "enabled" float exceptions by unmasking the exception masks in
MXCSR. This in turn had Dark Souls III's `expf` function to fault out,
as it checks if the MXCSR exception masks are set or not for determining
if underflow should assert or not.
Wow64/arm64ec has a similar problem where it always sets back to default
on signal. Which means game lose DAZ, but I'm not fixing that bug right
now.
Fixes#5391
ELF headers were read unconditionally because doing so was assumed to be cheap
(as the guest app would read them anyway shortly after). However, relocation
parsing was added since then, which has less predictable performance due to
crossing page boundaries and reading larger amounts of memory. It might be
possible to make the underlying code more efficient, but until that's done
it's better to skip this logic unless needed.
Closes#5390
We needed this handling on old kernels that didn't understand the
NOREPLACE flag. We no longer support kernels this old, so remove some of
this vestigial code.
Disables THP on some key locations that are fairly sparse
- rpmalloc
- This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
- These get in the hundreds of megabytes, while not being sparse they
trend towards only using a handful of pages and ballooning to 2MB
per thread is quite heavy.
- Lookup cache
- L1 specifically gets hit here which adds a decent chunk of overhead
due to sparsity.
Win32 for all of these also aren't handled, but that will need to be a
followup.
Add a no-op for zero length in ChangeProtectionFlags.
This fixes AMD Vivado 2025.2 which tries to mprotect with 0 as size and merge strategies fails:
```
Unexpected ChangeProtectionFlags Merge strategy! [0x400000, 0x401000) Versus [0x0, 0x0)
```
We were passing through 32-bit signed values to the 64-bit handler,
which isn't valid and we need to make sure to sign extend.
I don't think this fixes anything because negative values aren't valid,
but a negative could have been interpreted as a large positive without
this.
So WTF can see their usage. Previously didn't add this due to some weird
madvise bug that delete anon names, but that seems to be fixed in a
newer kernel.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.
In Bayonetta's title screen it went from 963MB down to 834MB resident.
When an anonymous FD is passed to execveat that has the CLOEXEC flag
set, then binfmt_misc fails with ENOENT.
Workaround this limitation by duplicating the FD, stripping its CLOEXEC
flag in the process.
For #5234
On 36-bit VA systems the stack was ending up /wherever/ when it should
be at the top of the VA space (usually).
Additionally VDSO was getting mapped anywhere on 64-bit, so push that to
the top of the VA space as well.
Also removes a check for old kernels not supporting MAP_FIXED_NOREPLACE.