logoalt Hacker News

abixbtoday at 8:24 PM5 repliesview on HN

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.


Replies

driverdantoday at 8:31 PM

> If this is truly AGI (subject to one's definition of AGI still)

Scoring well in a benchmark that's called AGI does not make an LLM AGI.

show 2 replies
anvuongtoday at 8:47 PM

It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.

thomasahletoday at 9:49 PM

• 98.6% on ARC-AGI-3

• 97.6% on frontier math

• 95.9% on CAD

• 100% on ExploitBench

Nothing modest about it

show 1 reply
mullingitovertoday at 8:31 PM

> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.

Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc

I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

show 4 replies
catigulatoday at 8:28 PM

They’re really, really scared because of the Mythos controversy. Skynet will be under hyped.