logoalt Hacker News

sailingparrotyesterday at 5:54 PM1 replyview on HN

Quite a jump in conclusion you are making here.

Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6.

[1]: https://x.com/aaron_lou/status/2079218392452530249


Replies

tristanjyesterday at 7:39 PM

My Fable 6 theory is admittedly speculative, but your “[public GPT-5.6] Sol is able to find the same counterexample” is also a jump in conclusions. Aaron specifically says he used “an internal version of Codex”. When asked whether that meant a different model or harness, he dodged the question, and only said the harness should be the standard commercial GPT-ultra harness [0]. He (intentionally) avoids identifying which model was used, so your claim is similarly unresolved. Given that Aaron works at OpenAI and has access to internal models, that GPT-5.7 is expected to launch in a few weeks and is rumored to be 10T+ parameters, it's very plausible that Aaron used that model in his analysis.

Furthermore, public GPT-5.6 pro failed six times to find a disproof to the Jacobian Conjecture [1], even with hints, which is evidence against the claim that the public GPT-5.6 Sol can solve this.

Regarding the existence of Fable 6: an internal upgraded version of Fable or Mythos almost certainly exists, given that Anthropic has been testing Mythos internally since April and previously released new models roughly every ~6 weeks.

[0] https://x.com/eliebakouch/status/2079237073001730510

[1] https://x.com/Tomodovodoo/status/2079172223055863895

show 1 reply