> I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
Isn't it fairly established that (generally [0]) manually written / optimized skill files perform a lot better than generated ones? Meaning that yes, this likely does hold.
[0] or to be specific, that the pecking order is: ai generated < human co/written < hyperoptimized for the specific model via some convergence process