logoalt Hacker News

intenextoday at 8:31 PM28 repliesview on HN

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.

Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.

I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.

For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.


Replies

mvkeltoday at 8:34 PM

Take it from the mouth of the creator of ARC-AGI:

When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.

show 6 replies
dingdong2026today at 10:03 PM

Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI.

Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor.

And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover.

At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life.

But sure, they can create a decent website or CRUD app, so they must be really smart.

That's AGI for you.

show 1 reply
zug_zugtoday at 9:16 PM

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).

Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):

- write a well-received book, write a best-seller

- come up with a new company idea, Run that company

- actually have a decent conversation, maybe someday talk somebody out of suicide effectively

- come up with its own ideas or theories that nobody else has presented

- understand the stock market well enough to trade better than an index fund

- come up with a theory of what makes games fun, make a popular game

- be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty of mass abundance and be able to argue persuasively)

- be able to articulate what it knows, what it doesn't know, and how what information it would need to bridge that gap

uludagtoday at 8:45 PM

Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?

Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.

show 1 reply
goochphdtoday at 8:53 PM

Small comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness.

[1] https://arcprize.org/blog/astra

vlmutolotoday at 9:40 PM

The ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design.

The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.

https://openai.com/index/how-two-settings-tripled-our-arc-ag...

morningbrewtoday at 8:37 PM

If your definition of AGI involves copy/pasting code and doing well in some made up benchmark, then probably AGI is close

abixbtoday at 8:32 PM

It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.

Eliezertoday at 9:49 PM

> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks

If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.

irthomasthomastoday at 9:42 PM

A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?

peratoday at 9:20 PM

To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.

Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.

show 2 replies
eggnettoday at 9:55 PM

AI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?

show 1 reply
bendergarciatoday at 9:51 PM

You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai

mbestotoday at 9:32 PM

> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

Simple. AGI is undefinable and benchmarks are notoriously flawed.

10xDevtoday at 9:41 PM

It will be AGI once it can update its own weights. It can't be "general" intelligence if its weights are frozen and requires to be updated manually.

adan1719today at 9:20 PM

The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI).

If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.

jbrittontoday at 9:43 PM

Watch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3

Then realize LLMs have zero of what anyone would consider intelligence.

irthomasthomastoday at 9:16 PM

ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.

regularfrytoday at 9:10 PM

In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.

dom96today at 9:32 PM

AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.

show 1 reply
qsorttoday at 8:58 PM

> The ARC-AGI-3 scorecard is extremely misleading (...)

True.

> Regardless, the result is still valid (...)

If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.

> in the sense of passing the most famous benchmark designed specifically to measure AGI progress

The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.

On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.

This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.

hypfertoday at 8:58 PM

What does "AGI" or "effective AGI" even mean, and why should anyone even care whether this unclear thing has been "reached" or not?

Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"?

I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.

show 1 reply
techpressiontoday at 9:45 PM

It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.

applfanboysbgontoday at 8:43 PM

My definition of AGI certainly doesn't entail passing a benchmark that some random person arbitrarily labelled AGI to make it sound cooler.

show 1 reply
yoz-ytoday at 9:08 PM

At this point? I’d like it to pass the Turing test and catch you in obvious lies. Not answering “no” to “can you hear me”.

It being able to comfortably say “i don’t know how to do this” rather than boiling and ocean to pick a shell from the shore without getting wet.

voidmain0001today at 9:12 PM

Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?

simianwordstoday at 8:45 PM

> where I am reasonably confident that there's essentially nothing that I am better than Fable

No. Humans are still better at super long context learning. Once that is beat you are completely correct.

show 1 reply
intrasighttoday at 8:32 PM

It has to pass the Turing test

show 2 replies