If the user passes in an absolute path then check to see if it exists in
the rootfs before executing.
Useful for launching applications directly out of the rootfs with
FEXInterpreter.
In the case that the absolute path doesn't exist in the rootfs then
fallback to the host system as usual
Adds a header only include utility folder that can be included from
everywhere.
Contains syscall helpers for older glibc and defines for older Linux
uapi headers missing some defines.
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.
This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.
Two minor behaviour changes that got mixed up with this change.
The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.
Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
Adds a new OutputSocket config option that when set will force all logs
to go through a socket.
This allows us to avoid polluting the guest application's output through
a socket instead of a file. Alleviating the issue of a file output
overwriting when multiple applications are ran.
Instead of early exiting, allow the application to continue running but
throw error messages anyway. Should allow some users to still run FEX
even if an application steals a page in the lower 32-bits
ERROR_AND_DIE was using __builtin_trap which would send our application
either a SIGILL or SIGTRAP depending on architecture.
This would then be captured by our faulting system and passed over to
the guest application.
If the guest application happened to have a signal handler installed for
these then it would pick up this fault and potentially continue
unsafely.
Now we can remove this usage of __builtin_trap and switch over to our
own handler.
Our frontend will check to see if the fault came from our handler and
uninstall the host signal handlers in this case. Which is what we want
for "ERROR_AND_DIE"
While we're in the area, lets move logging over to fmt to get it out of
the way.
This also gives us an opportunity to merge some CMakeLists things
together to organize it.
Previous we were just using an address hint to emulate MAP_32BIT.
Seemingly this behaviour has changed on AArch64 where it now isn't
guaranteed to scan up from the hint provided if exact allocation fails.
Now we pull in the full 32-bit allocator and add support for MAP_32BIT
in it. This limits the allocations there in to the first 2GB which Linux
expects.
Necessary for Mono's trampolines to work since it requires code to be in
the first 2GB on x86-64.
Only a partial fix for #1330, still needs preemption disabled to work.
On x86-64 hosts the Linux kernel resides in the top bit of VA which
isn't mapped in to userspace.
This means that userspace will never receive pointers living with that
top bit set unless you're running a 57bit VA host.
This results in userspace pointers never needing to do the sign
extending pointer canonicalization. But additionally some applications
actually don't understand the pointer canonicalization.
This results in bugs like: https://github.com/golang/go/issues/49405
Now if you're running on a 57bit VA host, this will end up behaving like
FEX but it seems like no one in golang land has really messed with 57bit
VA yet.
In AArch64, when configured with a 48bit VA, the userspace gets the full
48bit VA space and on EL mode switch has the full address range change
to the kernel's 48bit VA.
This means that we will /very/ likely allocate pointers in the high
48bit space since Linux currently allocates top-down.
So behave more like x86-64, hide the top 128TB of memory space from the
guest before boot.
Testing: Took the M1Max 15ms to 21ms allocate the top 128TB.
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
Puts the visibility of the main layer, application layers, and
environment in to FEXCore instead of FEX.
These layers aren't specific to FEX/FEXLoader and should live in
FEXCore.
Only the EmptyMapper remains in FEX, which should eventually move over
to FEXConfig since that is the only user.
Only the interpreter can run when executed as FD.
This happens when executed with binfmt_misc and can resolve an issue if
someone sets up the hardlink incorrectly.
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.
So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.
Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.
This information is only ever going to be offline. Will be useful for multiple reasons.
1) Searching for split lock usage in applications, which can be a programming bug.
a) This isn't visible on AMD systems and on Intel is a fairly new linux feature
2) Having more information about when an application breaks.
3) Useful for some minor profiling for devs looking for statistical data
Instead of watching to ensure our parent process is still alive. Mount the
squashfs once and use a combination of file leases and inotify to
ref count how many processes are using the rootfs.
This makes it so in the common case, the FEXMountDaemon only ever executes once
and runs until all FEX processes stop running.
In the rare edge case there is a race condition where multiple FEXMountDaemon
applications will start, but only one will end up mounting the squashfs.
In this case, one application wins and the one that failed to grab the lease
will wait until the other one completes.
With this change, squashfs should be reasonable to use now.
In the case that there is an immediate configuration failure. Use stderr specifically for outputting.
These errors won't be output typically because silent logging is enabled by default.
In the case of executable missing or rootfs configuration failure, print directly to stderr.
Previously it looked like FEX just exited for no reason.
We had multiple users encounter this and be confused
Only the frontends need to deal with ELF files specifically.
The backend doesn't need to be aware of them at all.
Since the ELF handling is the frontend's responsibility, move all the code to the frontend.
Due to our current allocation strategy. This variable was ending up in a weird state
where jemalloc allocated it using the glibc allocator.
On shutdown this was causing it to try and deallocate through FEX's allocator...Which
is very broken.
This lets linux clean up in this case, at least until the allocator lines are more
strongly written