My personal opinion is that they cheated evaluations to be able to release this pretending that it has no meaningful impact.
Otherwise, I don't see any logical explanation that some of their benchmark results would be higher when watermarked. Except if benchmark results are so unstable that they are an useless metric.
I agree that if they're so noisy they need to average over more runs, but maybe they didn't feel the need to bother.
As long as variance is nonzero, you won't get the exact same result twice for the same benchmark, so one of the numbers has to be higher and the other lower. If watermarking has no effect on the distribution (by construction, it should have no effect), that's a 50% chance the watermarked model gets the higher number. In this case, it happened 5 out of 8 times, which is hardly unusual.