It's a bit misleading I think because these benchmarks are for Max level, at which Anthropic newest models use crazy amount of reasoning tokens. And we know that intelligence scales with their number.