This flag breaks FEX heavily for now.
glibc 2.38 started using this flag as an optimization for posix_spawn.
It will fall back to a "non-optimized" implementation if the clone
syscall returns EINVAL. For now do this while we investigate a more
proper implementation.
Should be backported to 2312.1.
Split from #3284 without changing ownership semantics while I reduce the
debugging surface here.
Removes one usage of ParentThread from FEXCore. Which can be done since
it is no longer an opaque structure, we can read the StatusCode
directly.
No functional change.
Since we're invoking curl directly, we don't need to wrap it in `sh -c`
with this function.
Fixes an issue where curl downloads to a non-escaped path weren't
working. Now they do.
As I was poking around erofs-utils documentation, I found out that
fsck.erofs actually provides an option for extracting erofs images
without using fuse.
This finally puts the erofs handling on feature parity with squashfs.
O_TMPFILE has a few minor problems that I have been thinking about for a
while. I just recently got reminded about this and remembered that most
problems get resolved by using memfd_create.
- O_TMPFILE is only supported on some filesystems.
- Supported filesystem must be the one mounted to the pathname being
opened.
- Only a minor inconvenience as tmpfs and all related filesystems
support this.
- An inode is actually created on whatever filesystem is backing the
folder.
- `/tmp/` must exist as a directory
- If this folder happened to not be mounted then these temporary
files wouldn't have been created.
- memfd_create doesn't have a folder that needs to exist.
- We were leaving the files open as read/write
- While we were rewinding the file offset, an misbehaving application
could have wrote garbage to the temp file.
- memfd sealing allows us to open the FD as RW and then seal its
capabilities, making it a read-only FD.
- We were leaking FDs opened with O_CLOEXEC
- We could have just opened the O_TMPFILE with O_CLOEXEC
- memfd also just supports this flag, so use it.
- No real issues, just nice to be sanitary here.
Overall this doesn't really change any behaviour, but it is nice to
cleanup some of the edges there.
These are only used by gdbserver for filling out its XML data structures
so just remove them from FEXCore.
Also fixes the ordering on RegNames to match the definition of the enum
class definition in CoreState. This has been out of correct order since
we reordered registers months ago.
As we are moving more and more OS specific code to the frontend, this is
another set of functions that can be moved to FEXLoader from FEXCore.
No functional change here, only code moved from protected to private and
to FEXLoader's SignalDelegator.
Once more thread handling is moved to the frontend we can move even more
out of FEXCore. As follows:
- CheckXIDHandler can get moved.
- First pthread FEX makes would just call this.
- Register/UnregisterTLSState
- This can happen in the clone/thread handler once the frontend
handles it.
This leaves very little in the backend and is mostly an interface for
passing signal data to the frontend that it needs once a signal has
occured.
It additionally also is used for `SignalThread`.
The frontend needs to be in control of how threads are created. This is
inherent to the fact that OS threads are OS specific. We currently have
this weird split that when initializing the FEXCore context, we create a
parent thread at all times.
This does some initial cleanup that gets the core initialization nearly
decoupled.
Stop lying to the application about getcpu, sched_getaffinity, and
sched_setaffinity.
- getcpu would wrap the cpu result modulo the count of cores
- sched_setaffinity wouldn't work at all
- sched_getaffinity lied and always reported full affinity of config
option
Split off from #3282 to reduce burden.
We can read the data member directly now since it isn't opaque. In fact
we already do in the signal handlers. Removes these redundant helpers.
Removes one usage of ParentThread in FEXCore.
Suspend may be called on a thread before it has finished WOW64 initialisation,
keep track of all initialized threads and fallback to direct
NtSuspendThread when this is the case.
The frontend shouldn't need to know any information about how to
reconstruct eflags. Just give us the information we need and it'll work
out.
There are still some inherit limitations of this and some edge cases
that might give invalid data, but it is roughly as close as it was
before.
Just provide if the PC was in the JIT, the host GPRs, and the PState object from the signal
information and FEXCore does the rest.
We don't need to change the signature for `SetFlagsFromCompactedEFLAGS`
because during reloading of register state automatically does this for
us.
Many flag-generating instructions like cmp need to save calculations for
deferred PF and AF flag calculation. Currently, they require a store per flag,
which is prohibitively expensive for hot instructions like cmp. By instead
pinning PF/AF temporary results to registers (x26/x27 by convention here), we
eliminate many stores altogether and turn the rest into zero-cycle moves (on
64-bit at least, this isn't optimal for 32-bit emulation due to CTX->GetGPRSize
shenanigans, need to check if this requirement can be lifted..).
To implement, we model as SRA and then the existing SRA code is able to generate
good code with little manual tuning. (Future work will get us to excellent code
with more tuning ;) ).
The tradeoff is reducing the working dynamic GPR set by 2 registers, which might
increase spilling in some cases. I think it's worth it in practice, though.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Instruction count CI has transformed the way we work on FEX… I love the system
and want to make it better. there’s one part of instruction count CI that isn’t
so lovable: the problematic “optimal” flag on instructions.
There are several issues with this flag, both philosophical and practical.
– it is tedious to update the optimal flag when making an implementation
optimal. The effect of that is discouraging people from making instructions,
optimal, or encouraging people to fail to update the flag, and dilute the value
of it. Either way, since we care far more about optimal implementations, then we
do about updating the flag, clearly we should prioritize the implementation and
not the flag. This issue was not obvious at the outset, when instruction count,
CI was introduced, and still quite small. The problem magnified when we started
duplicating instructions in bulk for different combinations of CPU features
(flagm, AFP, etc.) that intern multiplies the manual work required to update the
flags by the corresponding constant factor. if it comes down to a choice between
removing this extra coverage and removing the flag, I think we all agree that
removing the flag is the lesser evil.
– The definition of “optimal” is fundamentally problematic. I have often
improved the instruction count of an instruction that was already “optimal”.
This is all kinds of silly, and calls into question whether there’s any value
whatsoever in the existing classifications of the flag. Furthermore, it is often
unknowable, whether an implementation really is optimal. Is it possible to
implement BZHI (with flag calculations) in fewer than eight instructions? We
don’t know, and it’s silly to pretend that we do.
– as a consequence of the problematic definitions , there are so many errors in
both directions that I don’t think there’s much value in preserving the existing
classification at the expense of +progress. Being able to say “32% of
instructions are translated optimally” is neat, but it really doesn’t tell us
anything whatsoever when you dig a little deeper.
So, as the flag is misleading at best and perhaps harmful at worst, let’s remove
it and make the instruction count CI, more useful overall. let’s let the
expected count and the assembly speak for themselves, and cut away the chaff. if
we want a meaningless number to report to management, we can instead calculate
the average blowup factor ;-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If a bind fails then this is usually because PID reuse happened and the
unix domain socket path still existed on the filesystem. Alternatively
the process did an execve which left a dangling unix domain file.
Unlink the file and try again in these cases. It will almost certainly
work the second time.
Instead of relying on sleeping in accept, use poll. This allows us to
early shutdown instead of leaving a thread hanging stuck in accept.
Doesn't change behaviour with FEX gdbserver waiting for attach, but in
the future when we all full process trees to have gdbserver running this
will be more important.
Using the server mount folder works most of the time, but when running
under pressure-vessel this stacks directories in a weird way because the
mount folder has some tricks applied to it.
Expose the temp folder being used directly instead.
Resolving the issue that we can only ever have one gdbserver process
running consuming port 8086 (even if the port number is cute).
Doesn't give us anything yet but in the future will allow us to have
whole process trees running gdbservers that we can attach to.
This fixes an issue where CPU tunables were ending up in the thunk
generator which means if your CPU doesn't support all the features on
the *Builder* then it would crash with SIGILL. This was happening with
Canonical's runners because they typically only support ARMv8.2 but we
are compiling packages to run on ARMv8.4 devices.
cc: FEX-2311.1
Requires #3249 to be merged first
Library alerting has been disabled for now, and storing IR while
gdbserver is running is removed.
Otherwise no functional change.
This lets most of the ASM tests run on 16K Linux hosts which is good because I
have a Mac and I'm bad at computer.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
When a syscall from the *at series is provided an FD but the path is
absolute then dirfd should be ignored. We weren't correctly doing this.
Now if the path is absolute, but set the argument to the special
AT_FDCWD..
Fixes#3204
There are some cases where we want to test multiple instructions where
we can do optimizations that would overwise be hard to see.
eg:
```asm
; Can be optimized to a single stp
push eax
push ebx
; Can remove half of the copy since we know the direction
cld
rep movsb
; Can remove a redundant insert
addss xmm0, xmm1
addss xmm0, xmm2
```
This lets us have arbitrary sized code in instruction count CI, with the
original json key becoming only a label if the instruction array is
provided.
There are still some major limitations to this, instructions that
generate side-effects might have "garbage" after the end of the block
that isn't correctly accounted for. So care must be taken.
Example in the json
```json
"push ax, bx": {
"ExpectedInstructionCount": 4,
"Optimal": "No",
"Comment": "0x50",
"x86Insts": [
"push ax",
"push bx"
],
"ExpectedArm64ASM": [
"uxth w20, w4",
"strh w20, [x8, #-2]!",
"uxth w20, w7",
"strh w20, [x8, #-2]!"
]
}
```
This allows us to use reciprocal instructions which matches precision of
what x86 expects rather than converting everything to float divides.
Currently no hardware supports this, and even the upcoming X4/A720/A520
won't support it, but it was trivial to implement so wire it up.