logoalt Hacker News

tristanjtoday at 6:45 PM6 repliesview on HN

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

Performance is significantly higher than Fable 5.1

Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/


Replies

scrlktoday at 6:47 PM

Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)

show 4 replies
andxortoday at 7:56 PM

> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

show 3 replies
leumontoday at 6:47 PM

The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.

With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

show 1 reply
opus5_hatertoday at 7:21 PM

any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.

show 3 replies
jjicetoday at 6:49 PM

100% on ExploitBench seems fitting given recent events.

malshetoday at 6:49 PM

I think we need a few writing related benchmarks.