The poster works at Anthropic, so they likely have internal access to the next generation of Fable. Their internal model is probably an absolute beast at mathematics, and the upcoming benchmark results will likely set a new record for maths performance.
I suspect this is what happened, because the poster is coy about sharing the actual prompt / reasoning trace used to reach this result. That would be covered by an NDA until the model is properly released.
Exciting times!
I'm not very excited. Access to the best AI is not a party I was invited to. And the people who are at that party, well, they don't exactly reflect on my best interests.
Quite a jump in conclusion you are making here.
Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6.
[1]: https://x.com/aaron_lou/status/2079218392452530249