glibc lazily initializes the SETXID signal handler until first thread
creation.
Once we create our first pthread, steal it back from GLIBC after the
fact.
Fixes SOMA again.
If we have SVE2 support and a vector length of at least 256, then we can
adequately handle AVX.
However we also add in the ability for application profiles to disable
AVX if necessary for any reason.
Not currently used, but will be in subsequent changes.
If an application is doing long jumps on exceptions then we need to
ensure that RIP is synchronized at least to block entry for some amount
of safety.
Due to our block linking which doesn't ensure that RIP is synchronized
on block entry, our exception handling wouldn't see a RIP change, thus
not going down the path that we jump to Dispatcher loop top on RIP
change.
Burn a couple of instructions on block entry to ensure that RIP
synchronized but only on config.
This does the setup for handling the named region object loading and
closing using the async interface.
This exercises the async interface while the async thread itself only
does the minimum no-op steps required to fake loading and saving.
The no-op interface is hooked up to the point of exercising it in the
most minimal of sense.
If the configuration is set to enable read-only or read/write object
code then it will spin up the async worker thread as well, but it
doesn't do anything yet.
Creates a pool allocator for OpcodeDispatcher and IRCompaction that
shares memory allocations between threads in a pool and supports
reclaiming stale allocations from participating threads.
A thread will use a heuristic to keep its claimed memory allocation
around if it is allocating a lot of code. If it slows down then it will
start putting the memory allocation back in to the thread pool.
Additionally if the allocation has been "disowned" and gone to sleep
while still retaining the allocation, then another thread can inspect
these stale allocations and reclaim it from the idling thread. Saving
further memory.
This needs some more work and cleanup but this is an interesting concept
that saves a decent amount of memory even in a basic test.
Causes teeworlds' title screen to go from 754MB to 599MB in my simple
test. 79.4% the memory usage is a good start.
Since we are masking signals before compiling code, we no longer will
receive a signal in the middle of compiling code.
This makes the compile service never be invoked so we can just remove
it.
We still have some locations in the syscall handling that isn't signal
safe, but compileservice wouldn't have fixed those anyway.
Had some idle time so I implemented this logic.
We do some tricky logic to have a big.little configuration even with
unknown CPU core types. Promoting or demoting a single MIDR depending on
if we have a mixed configuration or not.
In a non-hybrid design we only claim product names inside the CPUID
product string.
This will appear if you `/proc/cpuinfo` or read the CPUID registers
directly
eg on Snapdragon 888:
processor : 0
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 1
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 2
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 3
model name : FEX-2112-1-g13b14b85 Cortex-A55
processor : 4
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 5
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 6
model name : FEX-2112-1-g13b14b85 Cortex-A78
processor : 7
model name : FEX-2112-1-g13b14b85 Cortex-X1
eg on Macbook Pro VM which can't see the CPU type:
processor : 0
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 1
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 2
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 3
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 4
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 5
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 6
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
processor : 7
model name : FEX-2112-1-g13b14b85 Unknown ARM CPU
Several API functions act as state querying functions. These can take
some parameters by const to communicate that we don't intend to modify
the respective passed in instance.
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.
This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.
Two minor behaviour changes that got mixed up with this change.
The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.
Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
We can share this between the interpreter and the JIT. Necessary to
support the TSO-correct interpreter path.
With this change the interpreter is TSO-correct for GPRs. Just not FPRs
yet.
This lets us have JITsymbols grouped by library.
Useful for determining where to thunk.
Sadly perf doesn't have an option to deduplicate regions by name, so
some external tooling is necessary to make it look nice.
If the thunk configuration is enabled then preemptively load the
libraries. Since they will need to be loaded for any thunking library.
This is because we need to ALWAYS share the allocators to the thunks.
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
This was using the implicit thread TLS object. All users of this
have access to the thread object directly.
Use that instead. Fixes a subtle bug were the frontend could be trying
to do a callback and TLS sections weren't correctly set.
This always returns success and for easier state management, just return
our parent thread.
Frontend needs full visibility of this state anyway for thread
management.
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.
So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.
Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.