Benchmarking like that is often broken because of continuous CPU core clock speed adjustments, system interrupts, SMIs, etc.
I tried to fix it by switching hyperthreading off, playing with the scaling governor, boost, setting a CPU frequency to no avail. The jitter was too much and the results were not reproducible, so I just gave up.
Of course your mileage may vary; this was on an AMD Zen 3 CPU.
Used to do this sort of thing for computations that needed to run in the 10 microsecond range (HFT stuff), circa 2008. Had very predictable results because:
a) language was not garbage collected (C++)
b) we avoided heap lock contentions in critical paths by pre-allocating object pools at startup
c) I/O operations were offloaded to separate threads, connected by mutex locked linked lists
d) processing thread was bound to its own CPU core
That's about as deterministic as we could get.