pressure-vessel overrides our rootfs when it does a pivot_root.
Since we are still communicating to the FEXServer we were pulling the
configured rootfs.
Instead check if we are in pressure vessel and avoid doing that.
This fixes FEX running under pressure-vessel.
(There may be some implications to this down the road with code caching
but let's worry about that later)
Gets rid of some magic numbers and reduces the number of things that
need to manually change (e.g. when supporting AVX and needing to
increase the xmm size).
This class is similar to std::shared_mutex except it is safe for the
same thread to increment or decrement the ref counter multiple times.
This can be passed to regular std locks.
This will be required with the code object cache service soon.
Creates a pool allocator for OpcodeDispatcher and IRCompaction that
shares memory allocations between threads in a pool and supports
reclaiming stale allocations from participating threads.
A thread will use a heuristic to keep its claimed memory allocation
around if it is allocating a lot of code. If it slows down then it will
start putting the memory allocation back in to the thread pool.
Additionally if the allocation has been "disowned" and gone to sleep
while still retaining the allocation, then another thread can inspect
these stale allocations and reclaim it from the idling thread. Saving
further memory.
This needs some more work and cleanup but this is an interesting concept
that saves a decent amount of memory even in a basic test.
Causes teeworlds' title screen to go from 754MB to 599MB in my simple
test. 79.4% the memory usage is a good start.
Since we are masking signals before compiling code, we no longer will
receive a signal in the middle of compiling code.
This makes the compile service never be invoked so we can just remove
it.
We still have some locations in the syscall handling that isn't signal
safe, but compileservice wouldn't have fixed those anyway.
For 128-bit divides, we can very quickly check at runtime if we can
avoid the long divide and just do a 64-bit divide.
For unsigned just check if the top bits are all zero.
For signed just check if the top bits match bit 63 of the lower bits.
Additionally, keep the long divide handlers inside of the dispatcher.
This keeps the majority of the code bloat out of the code block itself,
significantly reducing block size for something doing these divides.
Also a fairly large icache improvement from this.
Hard performance number improvements here are hard to get since it
heavily depends on the application, also only occurs on x86-64.
Seems to have helped FTL and Dead Cells performance quite a bit though.
For the JIT cores we don't need to keep IR around, it's only necessary
for the Interpreter. So once the AOT IR service is done dealing with the
IR, check to see if we can delete it.
This causes teeworld's title screen memory usage to go from 730MB to
566MB. 77.5% the memory usage there.
This is effectively an infinite memory leak if the codespace wasn't ever
overwritten or invalidated. So larger memory usage programs would end up
having a larger impact.