Are you sure? I mean, yes of course tiny sample windows have this effect. But it seems at least possible that LLM is more prone to this effect when making estimates of human performance or behavior than when doing it for other topics.
In any case I feel the paper is interesting but almost begging to be misinterpreted.