While we're in the area, lets move logging over to fmt to get it out of
the way.
This also gives us an opportunity to merge some CMakeLists things
together to organize it.
Previous we were just using an address hint to emulate MAP_32BIT.
Seemingly this behaviour has changed on AArch64 where it now isn't
guaranteed to scan up from the hint provided if exact allocation fails.
Now we pull in the full 32-bit allocator and add support for MAP_32BIT
in it. This limits the allocations there in to the first 2GB which Linux
expects.
Necessary for Mono's trampolines to work since it requires code to be in
the first 2GB on x86-64.
Only a partial fix for #1330, still needs preemption disabled to work.
On x86-64 hosts the Linux kernel resides in the top bit of VA which
isn't mapped in to userspace.
This means that userspace will never receive pointers living with that
top bit set unless you're running a 57bit VA host.
This results in userspace pointers never needing to do the sign
extending pointer canonicalization. But additionally some applications
actually don't understand the pointer canonicalization.
This results in bugs like: https://github.com/golang/go/issues/49405
Now if you're running on a 57bit VA host, this will end up behaving like
FEX but it seems like no one in golang land has really messed with 57bit
VA yet.
In AArch64, when configured with a 48bit VA, the userspace gets the full
48bit VA space and on EL mode switch has the full address range change
to the kernel's 48bit VA.
This means that we will /very/ likely allocate pointers in the high
48bit space since Linux currently allocates top-down.
So behave more like x86-64, hide the top 128TB of memory space from the
guest before boot.
Testing: Took the M1Max 15ms to 21ms allocate the top 128TB.
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
Puts the visibility of the main layer, application layers, and
environment in to FEXCore instead of FEX.
These layers aren't specific to FEX/FEXLoader and should live in
FEXCore.
Only the EmptyMapper remains in FEX, which should eventually move over
to FEXConfig since that is the only user.
Only the interpreter can run when executed as FD.
This happens when executed with binfmt_misc and can resolve an issue if
someone sets up the hardlink incorrectly.
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.
So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.
Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.
This information is only ever going to be offline. Will be useful for multiple reasons.
1) Searching for split lock usage in applications, which can be a programming bug.
a) This isn't visible on AMD systems and on Intel is a fairly new linux feature
2) Having more information about when an application breaks.
3) Useful for some minor profiling for devs looking for statistical data
Instead of watching to ensure our parent process is still alive. Mount the
squashfs once and use a combination of file leases and inotify to
ref count how many processes are using the rootfs.
This makes it so in the common case, the FEXMountDaemon only ever executes once
and runs until all FEX processes stop running.
In the rare edge case there is a race condition where multiple FEXMountDaemon
applications will start, but only one will end up mounting the squashfs.
In this case, one application wins and the one that failed to grab the lease
will wait until the other one completes.
With this change, squashfs should be reasonable to use now.
In the case that there is an immediate configuration failure. Use stderr specifically for outputting.
These errors won't be output typically because silent logging is enabled by default.
In the case of executable missing or rootfs configuration failure, print directly to stderr.
Previously it looked like FEX just exited for no reason.
We had multiple users encounter this and be confused
Only the frontends need to deal with ELF files specifically.
The backend doesn't need to be aware of them at all.
Since the ELF handling is the frontend's responsibility, move all the code to the frontend.
Due to our current allocation strategy. This variable was ending up in a weird state
where jemalloc allocated it using the glibc allocator.
On shutdown this was causing it to try and deallocate through FEX's allocator...Which
is very broken.
This lets linux clean up in this case, at least until the allocator lines are more
strongly written