mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-08 13:00:18 +02:00
Due to how the 64-bit allocator previously worked, it was never subjected to memory regions larger than 64GB to be tracked. With the change in PR #1885, this has changed to have regions that will hit sizes larger than 170TB on some platforms. Better yet, even with smaller regions it still had a performance issue, it just wasn't as visible. First problem: We used MemSet instead of MemClear for the live page clearing. This caused pages to be claimed as "always in use". This would cause us to always scan the entire region on allocation, find that it didn't work and allocate a fresh region on every slab allocation. jemalloc saving us here since it allocates slabs from the OS fairly aggressively. Second problem: We used MemSet (now changed to MemClear) to "clear" the state tracking for pages. This causes ~600MB of memory to be used purely for state tracking. This was physically backed since we were writing to every bit of tracking for handling 256TB of VA. This had a fault dance with the kernel for every new page being hit here. Instead of clearing the the bits with a memset, clear it with madvise so it doesn't consume physical pages at all. This means we use significantly less physical memory for 32-bit applications. With this change, pressure-vessel startup time goes from 24 seconds down to 17 seconds. 70% of the original startup time. But really the main savings here comes from the memory reduction that PR #1885 ballooned, but has been an unseen problem before. Before that PR we were burning 2MB of physical memory per region for no reason. After that PR we were burning up to 600MB of physical memory per region for no reason. Changing a bit depending on how large the region ended up being. This now ends up being 2 pages starting out and grows as more pages are are used. A significant improvement.