While we were getting the application name for the application layer, we
were failing to store the filename for telemetry.
Save the filename we get for application layers and store it for the
telemetry file.
Otherwise these were just alway ending up as wine or wine-preloader.
In the case of an AArch64 builder is using 16kb or 64kb pages like is
common on servers then it would fail to compile, even if the resulting
application would only ever run on 4k page hosts.
Resolve this by removing the build check and hardcoding 4kb pages for
each of our uses. We still require 4kb pages to run, so this mostly just
removes the weirdness where it is 16kb builder + 4k runner. Would have
broken some of our assumptions when running.
This is just a memory leak waiting to happen.
Only the primary thread in an application really should have this set
since the kernel cleans it up.
We only ever allocate the primary thread of the guest application then
every host thread's stack on top of that. It's up to the guest when it
is cloning to set up new stack pointers, we don't manage that.
We are already allocating the first thread's size at the soft stack
limit with RLIMIT_STACK anyway.
Fixes#1556
Instead of just a basic mutex, also mask the signals.
This fixes the problem where we can end up receiving a signal in the
middle of memory allocation. Thus leaving the locked mutex in a broken
state.
This more closely matches the Linux kernel behaviour.
Since if you're in the middle of a memory allocating syscall, you won't
get signaled.
Adds a header only include utility folder that can be included from
everywhere.
Contains syscall helpers for older glibc and defines for older Linux
uapi headers missing some defines.
Migrates lingering instances of the old logger over to fmt where
applicable. This allows removing some of the old defines and functions.
The only remaining usages of the printf-based variant of the logger is
in Tests/LinuxSyscalls/Syscalls.cpp for the strace handling.
The CompileService was spinning up with the incoming thread mask and
then setting the mask once running.
Instead set the mask, which the thread inherits, then set it back once
it is created.
ERROR_AND_DIE was using __builtin_trap which would send our application
either a SIGILL or SIGTRAP depending on architecture.
This would then be captured by our faulting system and passed over to
the guest application.
If the guest application happened to have a signal handler installed for
these then it would pick up this fault and potentially continue
unsafely.
Now we can remove this usage of __builtin_trap and switch over to our
own handler.
Our frontend will check to see if the fault came from our handler and
uninstall the host signal handlers in this case. Which is what we want
for "ERROR_AND_DIE"
Only a partial fix for #1330, still needs preemption disabled to work.
On x86-64 hosts the Linux kernel resides in the top bit of VA which
isn't mapped in to userspace.
This means that userspace will never receive pointers living with that
top bit set unless you're running a 57bit VA host.
This results in userspace pointers never needing to do the sign
extending pointer canonicalization. But additionally some applications
actually don't understand the pointer canonicalization.
This results in bugs like: https://github.com/golang/go/issues/49405
Now if you're running on a 57bit VA host, this will end up behaving like
FEX but it seems like no one in golang land has really messed with 57bit
VA yet.
In AArch64, when configured with a 48bit VA, the userspace gets the full
48bit VA space and on EL mode switch has the full address range change
to the kernel's 48bit VA.
This means that we will /very/ likely allocate pointers in the high
48bit space since Linux currently allocates top-down.
So behave more like x86-64, hide the top 128TB of memory space from the
guest before boot.
Testing: Took the M1Max 15ms to 21ms allocate the top 128TB.
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.
So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.
Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.
When we are running a 32-bit process we end up mixing VMA Region allocators which
can cause crashes on shutdown.
This is because if a statically initialized object allocates memory, and then frees that
memory in the atexit handler. There is a chance that if it had to reallocate memory during the
VMA allocator switch, that the atexit handler will try freeing memory from the FEX VMA region allocator
AFTER it has already been deallocated.
This can't be safely worked around with atexit handlers.
So until we resolve this issue, we HAVE to leak the allocator so it can safely clean up and then let the kernel
clean up after us
This information is only ever going to be offline. Will be useful for multiple reasons.
1) Searching for split lock usage in applications, which can be a programming bug.
a) This isn't visible on AMD systems and on Intel is a fairly new linux feature
2) Having more information about when an application breaks.
3) Useful for some minor profiling for devs looking for statistical data
These parameters are used as indexes into the tracked memory, so sizing
the index variable relative to the type being stored is kind of sketchy.
e.g. If the tracked types were uint8_t for example, we'd still want the
indexing parameters to be regularly sized so that we aren't implicitly
truncating values all the time when passing values to a uint8_t
parameter (and while all usages are currently using uint64_t as type
T, we may as well address this).
While we're at it, we can make both Get() and operator[] const member
functions, since they don't directly modify any underlying data, they
only read it.
Only the frontends need to deal with ELF files specifically.
The backend doesn't need to be aware of them at all.
Since the ELF handling is the frontend's responsibility, move all the code to the frontend.
Transparent huge pages is a feature that the linux kernel opportunistically uses.
Depending on kernel configuration this feature is either enabled always, or when you madvise the region.
To ensure we hit both cases, madvise the regions we allocate in the 64-bit VMA allocator always
Can reduce kernel bookkeeping memory usage for our abusive allocator
After FEX has forked, there aren't any other threads in the process but their stacks remain.
We need to have some book keeping in place to have the stack ranges available to clean up
after fork.
We now keep both live stacks and dead stacks in a dequeue and on fork we will walk both to
clean up all stack objects that aren't our current thread.
Due to a disjoint mechanism inside of glibc we need to override both the glibc
publicly visible allocation functions AND the hooks.
We had broken the overriding when we changed the default visibility. So first fix that.
Then only override to FEX's allocators once we are ready for it.
Then also replace the hooks.
Additionally have a way to clear the hooks back to default.
Addresses #146 a little more by providing an interface to perform
fmt-compatible logging.
No more, will people on the project be tormented by classic printf
features like:
- Accidentally passing in a non-trivial type
- PRI macros
- Not being able to add support for custom types
- Mixing up signed/unsigned printf formatting specifiers accidentally
fmt-capable versions of the logging functions are named the same as the
existing functions, just with a "Fmt" or _FMT suffix (depending on
whether or not it's a function being used or a macro, respectively).
Overriding glibc's mmap and munmap functions don't encapsulate that functions they use
for allocating stack data so it was falling down the system mmap/munmap path.
This was causing host side stacks to end up in the 32-bit VA space in 32-bit applications.