In the case of nothing being set then with the default being zero now we
will not have fixed it up to calculate the number of threads based on
host core count.
Adds a header only include utility folder that can be included from
everywhere.
Contains syscall helpers for older glibc and defines for older Linux
uapi headers missing some defines.
Fixes thread, memory-map, and OS data packet types.
These were attempting to substr when the encode function already handles
that.
Was making it so gdb was only ever receiving the first 1000 bytes of the
data and then decoding incorrectly.
On x86 this is always supported.
On ARM this is only supported if FPCR writes actually enable the things.
Also detects the AFP feature for flushing input denormals to zero.
These are all part of the x86 MXCSR.
No Cortex supports FPCR exceptions, while Apple M1 CPUs support
Exceptions but not the true "AFP" extension
Apple instead supports some additional flags in their
`SYS_APL_AFPCR_EL0` register for enabling this.
If the CPU is unknown inside of the ARM CPU detection then the
MIDROption selected could have fallen down a path where it is set to
nullptr.
Resolve this crash by doing a nullptr check.
For the x86-64 JIT this is implemented with pulling rdtscp's result for
this value.
For Interpreter and AArch64 JIT this is implemented with the getcpu
syscall.
Theoretically AArch64 could implement this with MPIDR_EL1 but because
SoC vendors hecked this up, we can't. Thanks.
Kernel just returns zero + reserved bits if you try reading it.
Some of these were in the Emitter class and some were in the
HostFeatures.
Merge these together since in the future I'm going to be using all of
this data as a key for our AOT code cache.
Still only setting the new state if RIP is affected for now.
Noticed a bug where we weren't setting our frame RIP to the new RIP on
32-bit.
Decided to walk through more of the state setting while fixing that.
This gets #1214 further but then it eventually crashes with a read to
0x11.
This allows the compiler back these tables into the executable, which
reduces the amount of work the function has to do at runtime.
Reduces the runtime of this function by 30% relative to the previous commit.
The tables have recently been changed to be zeroed out as a whole on startup.
Speeds up InstallDebugInfo by about two orders of magnitude and reduces Debug
executable size by 1.8%.
Splits out the few required dependencies to a FEXCore_Base static
library.
FEXCore then links to this directly.
Then make it so the FEX Common code links to FEXCore_Base so the
jemalloc dependency doesn't get pulled in.
Had some idle time so I implemented this logic.
We do some tricky logic to have a big.little configuration even with
unknown CPU core types. Promoting or demoting a single MIDR depending on
if we have a mixed configuration or not.
In a non-hybrid design we only claim product names inside the CPUID
product string.
This will appear if you `/proc/cpuinfo` or read the CPUID registers
directly
eg on Snapdragon 888:
processor : 0
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 1
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 2
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 3
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 4
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 5
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 6
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 7
model name : FEX-2112-1-g13b14b85 Cortex-X1
eg on Macbook Pro VM which can't see the CPU type:
processor : 0
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 1
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 2
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 3
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 4
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 5
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 6
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 7
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
vixl has an assert check to ensure the register sizes are the same for
ubfm.
32-bit Inline syscalls hit this for both arguments and return value.
VCastFromGPR would have hit this but there aren't any x86 instructions
that move 8-bit and 16-bit values in to a vector register.
Several API functions act as state querying functions. These can take
some parameters by const to communicate that we don't intend to modify
the respective passed in instance.
It doesn't make any sense anymore to have specific instruction names
set for UND versus a nullptr string anymore.
Was useful when we could use it to determine the difference between
undefined from the start versus set in the tables but with unknown
decoding. Which is an edge case.
Now instead just zero initialize the data, which means it is an unknown
type and nullptr name. Which works for use.
Improves initialization time of the InitializeInfoTables function from
423 microseconds to 37 microseconds.
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.
This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.
Two minor behaviour changes that got mixed up with this change.
The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.
Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
This instruction zeroes a cacheline in memory that is weakly ordered and
non-temporal.
It uses the RAX register for where in memory to clear and aligns the
address on cacheline regardless of actual alignment.
This zeroes out an emulated 64byte cacheline. Writing zeros to memory.
This very specifically is only 64bytes to match x86 behaviour.
Also specifically non-temporal and weakly ordered. Which matches x86
CLZero behaviour.