logoalt Hacker News

softwaredougtoday at 6:40 PM4 repliesview on HN

I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)

https://venturebeat.com/technology/welcome-to-the-agi-era-op...


Replies

aabhaytoday at 7:26 PM

This is with the caveat that OpenAI uses their own harness for this:

> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

show 2 replies
kaspernitoday at 6:48 PM

"On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.

But the comparison isn't straightforward.

OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."

arctic-truetoday at 6:47 PM

The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)

show 1 reply
Bluesteintoday at 6:43 PM

100%, some say.-