logoalt Hacker News

GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?

67 pointsby mbaumantoday at 2:56 PM13 commentsview on HN

Comments

draginoltoday at 7:51 PM

So Fable "won" but it cost $124.76 for marginal performance benefits over the $22.56 5.6 Sol run.

xnorswaptoday at 5:13 PM

A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.

Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.

show 1 reply
giwooktoday at 4:06 PM

Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?

show 1 reply
hartatortoday at 4:47 PM

It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.

effnorwoodtoday at 7:18 PM

Define "best" and "performs"

grim_iotoday at 3:58 PM

I'd expect google to do well here, since they were historically strong at multimodal and physics.

show 1 reply
arisAlexistoday at 6:09 PM

Google with apptronic should have good models soon

jespineltoday at 3:57 PM

Nice! It is missing Codex in the agent harnesses comparison IMO.

gizmodo59today at 4:00 PM

Yet another "benchmark to promote their own harness"