logoalt Hacker News

the8472 • today at 8:42 AM • 0 replies • view on HN

Swap latency explodes when you get into a swap storm where tasks continuously touch different pages and fight to get things paged in and out, forming a queue waiting for swap IO. The other thing causing multi-second pauses is approaching OOM and the system attempting some last-ditch efforts before unleashing the OOM reaper, but that's not strictly swapping, it'd also happen on a swapless system.

If it's only a single fault I'd expect it to be to be dominated by IO latency, which is pretty good on NVMe. As the article shows those 40ms were accumulated over hundreds of pagefaults. The problem there was that there was a stop-the-world pause stalled by all those pagefaults together.

A fully-concurrent and swap-friendly GC you could maybe define that it only increases the active working-set by x GB on top of what the application itself (and the rest of the system) is actively using and will only page-in y GB per second. And it'd do so on GC threads, not the application threads. Whether that would be sufficient to not cause a swap storm would still depend on all the other stuff happening besides the GC.