A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.
Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.
Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?
It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.
Define "best" and "performs"
I'd expect google to do well here, since they were historically strong at multimodal and physics.
Google with apptronic should have good models soon
Nice! It is missing Codex in the agent harnesses comparison IMO.
Yet another "benchmark to promote their own harness"
So Fable "won" but it cost $124.76 for marginal performance benefits over the $22.56 5.6 Sol run.