LLMs don't search trees. They generate plausible proofs and a human has to check it's true. Repeat until.
That's not what happened here. This isn't a proof; it's a counterexample. The model was perfectly capable of verifying its correctness. You could have verified it by hand if you wanted; the verification is trivial. Finding it was the hard part.
Agents can generate formal proofs that are checked with an oracle like Lean and can run in a loop.
That's not what happened here. This isn't a proof; it's a counterexample. The model was perfectly capable of verifying its correctness. You could have verified it by hand if you wanted; the verification is trivial. Finding it was the hard part.