I don't think you can extrapolate that measurement across multiple sequential draws like that. We presumably are comparing against a single trajectory rather than a tree of trajectories. So once we make the wrong choice and step off of the blessed path, we have no way to assign a ranking to the next token; it's error is undefined.
I've seen LLMs self correct in chains of thought ("because of foo and bar, I need to... Wait, bar is not true, so that won't work") so I have to imagine this is a massive overestimate, errors do not necessarily compound.
actually, bar is true
but wait, the models constantly go back and forth on these things in their thinking traces, so it is unclear which self correcting is actually correct
I would say they do compound until proven otherwise.
Having "Wait, bar is not true, so that won't work" is not necessarily a correction. In fact, the problem is: across a long text it is a correction of a single mistake, but we are talking about thousands here.
But yes, of course that was a rough estimate. But the problem is - we don't really know what we are measuring here. Maybe there's a 2,000,000x difference of intelligence between coding indexes 52 and 50. By some measure that just feels small because that's how we process it akin to audio db.
Regardless the point is KLD and whatever they came up with is not meaningful. And they did not publish comparisons on real benchmarks.