Had some idle time so I implemented this logic.
We do some tricky logic to have a big.little configuration even with
unknown CPU core types. Promoting or demoting a single MIDR depending on
if we have a mixed configuration or not.
In a non-hybrid design we only claim product names inside the CPUID
product string.
This will appear if you `/proc/cpuinfo` or read the CPUID registers
directly
eg on Snapdragon 888:
processor : 0
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 1
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 2
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 3
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 4
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 5
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 6
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 7
model name : FEX-2112-1-g13b14b85 Cortex-X1
eg on Macbook Pro VM which can't see the CPU type:
processor : 0
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 1
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 2
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 3
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 4
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 5
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 6
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 7
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
Several API functions act as state querying functions. These can take
some parameters by const to communicate that we don't intend to modify
the respective passed in instance.
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.
This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.
Two minor behaviour changes that got mixed up with this change.
The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.
Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
This was using the implicit thread TLS object. All users of this
have access to the thread object directly.
Use that instead. Fixes a subtle bug were the frontend could be trying
to do a callback and TLS sections weren't correctly set.
This always returns success and for easier state management, just return
our parent thread.
Frontend needs full visibility of this state anyway for thread
management.
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.
So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.
Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.
I saw a red herring that I thought the high cpu usage in steamwebhelper could come from signal handlers.
This turned out to not be the case, but now I've got this implemented.
Installs the few signal handlers that we need upfront but for everything that isn't a mandatory signal
we instead now wait until the guest also installs that signal handler.
This fixes#1107
Leafs come from ECX but only some CPUID functions support this.
This adds the initial infrastructure but doesn't yet add support for the CPUID functions to consume the leaf.
Patch moves HandleSIGSEGV in HostFactory.cpp to a non-frontend host signal handler, then registers its own frontend signal handler to catch unhandled segfaults.
Sadly these things can't be split without breaking functionality so it
turns in to a bit of a mess.
SyscallHandler is very much something that is a Linux only construct and
shouldn't be in FEXCore itself. Lets the frontend register a
Syscallhandler with FEXCore. FEXCore itself is then aware of the current
syscall ABI and handles the ABI in an optimal fashion.
So it is not a 100% clean break otherwise we would lose performance.
The SignalDelegator then needs to move to the frontend since the
SyscallHandler requires it for signal based syscalls.
The CPU backend signal handling still needs to happen in FEXCore because
it is a very tight coupling with the CPU backend.
Once we need to support more Signal handling we can give the backends
cleaner support to select which specific OS handler to handle.
This interface is for native host code to be able to call back in to JIT
to execute guest code.
This is mainly for thunks but could be used for other purposes.
This can cross a couple ABI boundaries so one must be careful when
calling in to this
This requires implementing custom ASM dispatchers for the interpreter
side of things so it works correctly.
Allows all of our CPU backends to safely support signaling.
This is a bit of a nightmare change and requires rethinking logic about
debugging in some instances.
Splits out to a jump table approach and loosely classifies and groups
the syscalls.
Also defines syscalls by number of arguments so it can be optimized