Had some idle time so I implemented this logic.
We do some tricky logic to have a big.little configuration even with
unknown CPU core types. Promoting or demoting a single MIDR depending on
if we have a mixed configuration or not.
In a non-hybrid design we only claim product names inside the CPUID
product string.
This will appear if you `/proc/cpuinfo` or read the CPUID registers
directly
eg on Snapdragon 888:
processor : 0
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 1
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 2
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 3
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 4
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 5
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 6
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 7
model name : FEX-2112-1-g13b14b85 Cortex-X1
eg on Macbook Pro VM which can't see the CPU type:
processor : 0
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 1
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 2
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 3
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 4
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 5
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 6
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 7
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
Several API functions act as state querying functions. These can take
some parameters by const to communicate that we don't intend to modify
the respective passed in instance.
It doesn't make any sense anymore to have specific instruction names
set for UND versus a nullptr string anymore.
Was useful when we could use it to determine the difference between
undefined from the start versus set in the tables but with unknown
decoding. Which is an edge case.
Now instead just zero initialize the data, which means it is an unknown
type and nullptr name. Which works for use.
Improves initialization time of the InitializeInfoTables function from
423 microseconds to 37 microseconds.
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.
This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.
Two minor behaviour changes that got mixed up with this change.
The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.
Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
Without this, the internal BucketList type will always be allocated with
T as a uint32_t, due to the default type for T, even if T is specified
differently in other code.
In quite a few places we have a raw primitive to represent an IR node's
ID. This can make reading some bits of the API (or the passes) a little
confusing to take in, since there's no meaningful type name for some
data structure members.
We can provide an alias that communicates this directly to the reader.
This also has the nice benefit of providing a single point of definition
for node IDs which can allow for easier changes in the future (e.g.
making Node IDs strongly-typed etc).
Allows BucketList to work with types that aren't a direct numeric
primitive, so long as the object has equality operators defined and are
default constructible
Migrates lingering instances of the old logger over to fmt where
applicable. This allows removing some of the old defines and functions.
The only remaining usages of the printf-based variant of the logger is
in Tests/LinuxSyscalls/Syscalls.cpp for the strace handling.
The CompileService was spinning up with the incoming thread mask and
then setting the mask once running.
Instead set the mask, which the thread inherits, then set it back once
it is created.
If a SIGBUS is received in compile service code then we weren't handling
it correctly. Instead we would fail the JIT space check and hand it off
to the guest.
Fixes#1217
Instead of throwing an error and closing down FEX. Instead pass the
SIGILL to the guest application.
On unhandled instruction implementation the instruction, we instead emit
a _Break IR op at that location.
A _Break IR op will ensure the context state is synchronized at the
point of of the fault and has fairly low overhead. We branch to the
dispatcher which does the SRA spilling.
Tested this with an application that attempts an AVX512 instruction,
catches the fault, and continues onward.
With #1383 in place, we also won't pass spurious ERROR_AND_DIE to the guest anymore.
ERROR_AND_DIE was using __builtin_trap which would send our application
either a SIGILL or SIGTRAP depending on architecture.
This would then be captured by our faulting system and passed over to
the guest application.
If the guest application happened to have a signal handler installed for
these then it would pick up this fault and potentially continue
unsafely.
Now we can remove this usage of __builtin_trap and switch over to our
own handler.
Our frontend will check to see if the fault came from our handler and
uninstall the host signal handlers in this case. Which is what we want
for "ERROR_AND_DIE"
In a syscall microbench this improves performance by ~19% on my
Snapdragon 888.
Going from ~9.6 million syscalls per second to ~11.5 million.
Macbook Pro is less effective here due to high syscall overhead due to
VM. Going form 7.2M/s to 7.6M/s, ~6% improvement
We can also inline some 32-bit syscalls but that will need some more
work which isn't done yet. Even though the op in the JIT supports it.
Ran in to this when removing syscall nodes. If an IR op has side-effects
then this generic helper can not remove them.
First time this was encountered and it was confusing
These aren't strictly necessary for enum flags and through discussion in
\#1363, would lead to an awkward to use overload.
If these are ever needed, they can be added back at a later date.
Only a partial fix for #1330, still needs preemption disabled to work.
On x86-64 hosts the Linux kernel resides in the top bit of VA which
isn't mapped in to userspace.
This means that userspace will never receive pointers living with that
top bit set unless you're running a 57bit VA host.
This results in userspace pointers never needing to do the sign
extending pointer canonicalization. But additionally some applications
actually don't understand the pointer canonicalization.
This results in bugs like: https://github.com/golang/go/issues/49405
Now if you're running on a 57bit VA host, this will end up behaving like
FEX but it seems like no one in golang land has really messed with 57bit
VA yet.
In AArch64, when configured with a 48bit VA, the userspace gets the full
48bit VA space and on EL mode switch has the full address range change
to the kernel's 48bit VA.
This means that we will /very/ likely allocate pointers in the high
48bit space since Linux currently allocates top-down.
So behave more like x86-64, hide the top 128TB of memory space from the
guest before boot.
Testing: Took the M1Max 15ms to 21ms allocate the top 128TB.