logoalt Hacker News

scrlktoday at 6:47 PM4 repliesview on HN

Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)


Replies

tedsanderstoday at 7:48 PM

Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

(I coauthored the linked blog post)

woahtoday at 7:37 PM

Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?

show 1 reply
kaspernitoday at 6:51 PM

yes it is.

enraged_cameltoday at 7:37 PM

Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.

show 1 reply