Love it. Unfortunately, for some benchmarks, it can be bit difficult to get representative inputs which take hundreds of milliseconds. I'm curious as to why the author doesn't loop 1-10ms inputs to deal with variance? Which also deals with startup costs. Rust microbench harnesses were already good at this stuff-out of-the-box (run-to-run variance, etc) several years ago.