Fixes#2136
This is a fairly tricky edge case to support with FEX.
If execveat is used with AT_EMPTY_PATH then the application can pass an
FD to execve instead of a filename. This includes FDs that have been
deleted from the disk so the child process can't open it by filename
anymore.
To work around this limitation, we need to pass the FD to the new FEX
process and open it directly, similar to how binfmt_misc works with FDs.
The FD will get passed through environment variables, which the new
process will check for and then remove the variable from the
environment.
Lots of prickly edge cases to support here.
Without binfmt_misc:
- Passes the FD to FEXLoader directly.
- Requires duplicating the FD if it has O_CLOEXEC on the FD.
With binfmt_misc:
- Shebang file, pass directly to FEXLoader, just like without binfmt.
- x86 ELF Files, rely on the kernel's binfmt_misc support here.
- Unsupported ELF files, let kernel handle it through binfmt_misc
Argument handling:
- The application can pass in no arguments.
- Means our application configurations were failing to find a config
- Also various checks in the frontend were failing.
- If opened through an FD, find the symlink for that FD for the
application configuration instead.
Side note:
Fixed a performance issue in execve where when we were checking for file
format support. Either ELF or Shebang files, we were reading the /whole/
file upfront. We only need to read a header worth of ELF files, and only
257 bytes if it is potentially a shebang file. Should dramatically
reduce some application's execve times.
This is causing some heartburn with the hardlink.
- Removes some termux cmake list hacking.
- Removes the need to do post-install packaging fixups when hardlinks get dropped.
- Removes a custom uninstall target that was necessary before.
`std::filesystem::exists` is particularly gnarly in how it checks to see
if the file exists.
It allocates memory, it creates lists, it splits things, then eventually
checking the status
Remove all this overhead to help out minorly for applications that
execve a lot.
This creates a generic interface that FEXCore can use for timeline
profiling. This allows us to create a generic interface which the
backend details are hidden so we can support multiple timeline profile
APIs.
The only API supported right now is ftrace/gpuvis. Which is extremely
lightweight of an interface with minimal overhead.
We must be careful here since in most cases will will have dozens of
FEX instances running at any given time. So a timeline profiler like
Microprofiler can have major issues since that only ever expects a
single process at a time.
Not enabled by default but just needs the `ENABLE_FEXCORE_PROFILER`
cmake option set to enable.
Implements AT_PLATFORM: Ends up being `i686` or `x86_64` depending on
ELF arch
Implements AT_HWCAP and AT_HWCAP2
AT_HWCAP is just CPUID function 01h EDX result
AT_HWCAP2 only has two defined bits in it, which we don't support
either.
Implements AT_RANDOM
Previously we were just sticking hardcoded values in to this.
Now we pass along the host's AT_RANDOM, or we generate our own if that
doesn't exist
Fixes#788
Also in FEXLoader make sure to use `EraseSet` for these runtime options.
Fixes a bug where the config was being set to nothing, breaking the
ThunksDB configuration option.
We have separate configurations for the Application path versus the
application name we are using as a configuration choice.
Example 1: FEXBash "wine Crysis64.exe"
Previous APP_FILENAME will contain `/usr/bin/wine`, which is still used
elsewhere.
This new APP_CONFIG_NAME will contain `Crysis64.exe`
Example 2: FEXBash glxgears
Previous APP_FILENAME will contain `/usr/bin/glxgears`
APP_CONFIG_NAME will contain `glxgears`
We didn't have this exposed any other way before.
While we were getting the application name for the application layer, we
were failing to store the filename for telemetry.
Save the filename we get for application layers and store it for the
telemetry file.
Otherwise these were just alway ending up as wine or wine-preloader.
This is a relatively invasive change since multiple things needed to
happen at once.
* Socket based logging is removed
* Logging has been replaced to only support stdout, stderr, and server
* Server is now default and replaces what FEXLogServer did
* Server logging now uses a pipe instead of a socket
* Can be faster than stderr and stdout since the application doesn't
need to wait on terminal output
* FEXMountDaemon has been removed
* Functionality has been merged in to FEXServer
* FEXServer is always executed on FEX initialization time
* Similar in behaviour to Wine's wineserver
* Can explicitly start this before using FEX for logging purposes
* Stays around until all instances of FEX exit
* Will stick around for a short amount of time in case of spurious
execution
* FEXServer will soon be extended to do more than logging and squashfs
mounting
* FEX rootfs scripts will need to be updated to support this path
* Just means rbinding the /tmp folder and forcing a FEXServer instance
to be alive
* Pressure-vessel works fine in this case since FEXServer will already
be running
* It already rbinds the host /tmp folder which is why this works
FEX always builds with PIE and we don't hit this issue anymore anyway.
If some application wants to inject a page in to the lower 32-bits then
we have no reason to complain about it anymore. Just let it go and
hopefully they know what they are doing.
Fixes#1559
If the user passes in an absolute path then check to see if it exists in
the rootfs before executing.
Useful for launching applications directly out of the rootfs with
FEXInterpreter.
In the case that the absolute path doesn't exist in the rootfs then
fallback to the host system as usual
Adds a header only include utility folder that can be included from
everywhere.
Contains syscall helpers for older glibc and defines for older Linux
uapi headers missing some defines.
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.
This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.
Two minor behaviour changes that got mixed up with this change.
The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.
Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
Adds a new OutputSocket config option that when set will force all logs
to go through a socket.
This allows us to avoid polluting the guest application's output through
a socket instead of a file. Alleviating the issue of a file output
overwriting when multiple applications are ran.
Instead of early exiting, allow the application to continue running but
throw error messages anyway. Should allow some users to still run FEX
even if an application steals a page in the lower 32-bits
ERROR_AND_DIE was using __builtin_trap which would send our application
either a SIGILL or SIGTRAP depending on architecture.
This would then be captured by our faulting system and passed over to
the guest application.
If the guest application happened to have a signal handler installed for
these then it would pick up this fault and potentially continue
unsafely.
Now we can remove this usage of __builtin_trap and switch over to our
own handler.
Our frontend will check to see if the fault came from our handler and
uninstall the host signal handlers in this case. Which is what we want
for "ERROR_AND_DIE"
While we're in the area, lets move logging over to fmt to get it out of
the way.
This also gives us an opportunity to merge some CMakeLists things
together to organize it.
Previous we were just using an address hint to emulate MAP_32BIT.
Seemingly this behaviour has changed on AArch64 where it now isn't
guaranteed to scan up from the hint provided if exact allocation fails.
Now we pull in the full 32-bit allocator and add support for MAP_32BIT
in it. This limits the allocations there in to the first 2GB which Linux
expects.
Necessary for Mono's trampolines to work since it requires code to be in
the first 2GB on x86-64.
Only a partial fix for #1330, still needs preemption disabled to work.
On x86-64 hosts the Linux kernel resides in the top bit of VA which
isn't mapped in to userspace.
This means that userspace will never receive pointers living with that
top bit set unless you're running a 57bit VA host.
This results in userspace pointers never needing to do the sign
extending pointer canonicalization. But additionally some applications
actually don't understand the pointer canonicalization.
This results in bugs like: https://github.com/golang/go/issues/49405
Now if you're running on a 57bit VA host, this will end up behaving like
FEX but it seems like no one in golang land has really messed with 57bit
VA yet.
In AArch64, when configured with a 48bit VA, the userspace gets the full
48bit VA space and on EL mode switch has the full address range change
to the kernel's 48bit VA.
This means that we will /very/ likely allocate pointers in the high
48bit space since Linux currently allocates top-down.
So behave more like x86-64, hide the top 128TB of memory space from the
guest before boot.
Testing: Took the M1Max 15ms to 21ms allocate the top 128TB.
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
Puts the visibility of the main layer, application layers, and
environment in to FEXCore instead of FEX.
These layers aren't specific to FEX/FEXLoader and should live in
FEXCore.
Only the EmptyMapper remains in FEX, which should eventually move over
to FEXConfig since that is the only user.
Only the interpreter can run when executed as FD.
This happens when executed with binfmt_misc and can resolve an issue if
someone sets up the hardlink incorrectly.
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.
So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.
Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.
This information is only ever going to be offline. Will be useful for multiple reasons.
1) Searching for split lock usage in applications, which can be a programming bug.
a) This isn't visible on AMD systems and on Intel is a fairly new linux feature
2) Having more information about when an application breaks.
3) Useful for some minor profiling for devs looking for statistical data
Instead of watching to ensure our parent process is still alive. Mount the
squashfs once and use a combination of file leases and inotify to
ref count how many processes are using the rootfs.
This makes it so in the common case, the FEXMountDaemon only ever executes once
and runs until all FEX processes stop running.
In the rare edge case there is a race condition where multiple FEXMountDaemon
applications will start, but only one will end up mounting the squashfs.
In this case, one application wins and the one that failed to grab the lease
will wait until the other one completes.
With this change, squashfs should be reasonable to use now.