mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-10 23:00:30 +02:00
While tracking issues in #3162, I had encountered a random crash that I started hunting. It was very quickly apparent that this crash was unrelated to that PR. I just happened to be running a unittest that was creating and tearing down a bunch of threads that exacerbated the problem. See as follows with the strace output: ``` [pid 269497] munmap(0x7fffde1ff000, 16777216) = 0 [pid 269497] munmap(0x7fffde1ff000, 16777216 <unfinished ...> [pid 268982] mmap(NULL, 16777216, PROT_READ|PROT_WRITE|PROT_EXEC, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7fffde1ff000 [pid 269497] <... munmap resumed>) = 0 ``` One thread is freeing some memory with munmap, another one then does a mmap and gets the same address back. Nothing too crazy at initial glance, but taking a closelier look, we can see that there are two strange oddities: 1) We are double unmapping the same address range through munmap 2) The second munmap is interrupted and returns AFTER the mmap. This has the unfortunate side-effect that the mmap that just returned the same address has actually just been unmapped! This was resulting in spurious crashes around thread creation that was SUPER hard to nail down. The problem comes down to how code buffer objects are managed, in particular how the Arm64Emitter and Dispatcher handled its buffers. Arm64Emitter is inherited by two classes; Dispatcher, and Arm64JITCore. On class destruction the emitter would free its internal tracking buffer. Additionally on destruction, the Arm64JITCore would walk through all of its CodeBuffers and free them. The problem ends up being that in the Arm64JITCore, it would free its code buffers which also ended up being the current active buffer bound to the Arm64Emitter. Thus causing the Arm64Emitter to come back around and try to free the same buffer again. This is a double-free problem! and was only visible on thread exiting! Can't track double frees with mmap and munmap with current tooling! This problem typically didn't occur because of how fast the destruction usually takes and jemalloc inbetween also typically means the problem doesn't occur. Initially thinking this was a threaded pool allocator bug because typically the new allocation would end up in there once a new thread was spinning up. Now we change behaviour, Arm64Emitter doesn't do any buffer management itself, instead just passing an initial buffer on to its internal buffer tracking if given one up front. This leaves the Dispatcher and the Arm64JITCore to do their buffer management and ensuring there is no double free. The day is saved!