logoalt Hacker News

xnorswaptoday at 5:13 PM1 replyview on HN

A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.

Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.


Replies

mrr7337today at 5:43 PM

These benchmarks are done with incorrectly and missing a lot of baselines. The article seems very vibe coded too.