logoalt Hacker News

GPT-6 Astra

868 pointsby kibaetoday at 6:41 PM607 commentsview on HN

System Card: https://deploymentsafety.openai.com/gpt-6-astra

Related ongoing threads:

OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691

GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147


Comments

dangtoday at 7:37 PM

Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273

How about we stick to that one for talking about the rollout, and this one for talking about the model?

intenextoday at 8:31 PM

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.

Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.

I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.

For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

show 15 replies
Planktonnetoday at 8:32 PM

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.

It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.

This is farcical.

show 8 replies
abixbtoday at 8:24 PM

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.

show 4 replies
Chinjuttoday at 8:53 PM

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)

show 5 replies
dalemhurleytoday at 8:40 PM

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.

Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).

Codex is slightly better than Claude Code.

Good on Sam Altman getting back to basics and turning OpenAI around.

show 4 replies
manlymuppettoday at 9:27 PM

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?

Even if I did trust an AI to get everything right, it's not like the AI can read my mind.

If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?

All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.

(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)

tristanjtoday at 6:45 PM

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

Performance is significantly higher than Fable 5.1

Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

show 6 replies
HAL3000today at 8:00 PM

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.

I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.

Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.

Canceling my Anthropic Max sub when this ships.

show 2 replies
XCSmetoday at 8:58 PM

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

show 3 replies
astrobiasedtoday at 9:11 PM

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547

Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.

It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.

The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?

With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.

jdprgmtoday at 8:41 PM

Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.

It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.

I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.

show 8 replies
x312today at 7:50 PM

Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?

show 5 replies
Cu3PO42today at 7:40 PM

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.

[0] https://arxiv.org/abs/2608.31126

[1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

show 8 replies
maherbegtoday at 8:49 PM

Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?

maybe call it EngEmployeeBench

isoprophlextoday at 7:42 PM

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Well that sounds like fun. It has become better at hiding its thoughts.

show 7 replies
GodelNumberingtoday at 8:36 PM

The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:

Terminal-Bench 4.0: High (57.9%), Max (56.7%)

DeepSWE: High (73.3%), Max (71.5%)

It _loses_ 1-2% performance going to High from Max

show 1 reply
tintortoday at 7:47 PM

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

show 6 replies
udbhavstoday at 8:22 PM

I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.

show 1 reply
softwaredougtoday at 6:40 PM

I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)

https://venturebeat.com/technology/welcome-to-the-agi-era-op...

show 4 replies
rcr-antitoday at 8:45 PM

The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.

pandinustoday at 8:10 PM

Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.

putlaketoday at 7:49 PM

> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

Not on Azure? If so, that's a big deal.

show 4 replies
bmenrightoday at 9:27 PM

> GPT‑6 Astra brings together years of research and big bets across pre-training

Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?

show 1 reply
BeetleBtoday at 8:02 PM

It's been over an hour, Simon! Where's the Pelican?

show 1 reply
itissidtoday at 9:34 PM

All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.

Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.

swalshtoday at 7:37 PM

I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.

show 1 reply
petilontoday at 7:59 PM

This is wild: OpenAI is basically declaring that AGI is here.

https://www.theverge.com/ai-artificial-intelligence/989601/o...

“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”

show 12 replies
toshtoday at 6:44 PM

$10 per million input tokens and $50 per million output tokens

sol is $4 / $20

show 3 replies
wiseowisetoday at 9:29 PM

Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?

aliljettoday at 7:48 PM

The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...

show 2 replies
serjestertoday at 9:03 PM

Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.

theseamusjamestoday at 8:05 PM

Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.

show 1 reply
codruterdeitoday at 8:57 PM

I was actually wondering when they will release the new Opel Astra model. Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.

MASNeotoday at 7:54 PM

Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…

oh_notoday at 7:45 PM

Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.

alpinemantoday at 9:15 PM

“allowing non-technical people to create and play custom games that go beyond rudimentary elements”

Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

aliljettoday at 7:15 PM

The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?

show 3 replies
trixn86today at 8:18 PM

Secret tip to win the mario cart clone: Just hold w, no steering needed.

KronisLVtoday at 9:15 PM

It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.

_ache_today at 7:36 PM

https://ache.one/gpt6_now_down.png

Big claims, expensive and not release to the public yet.

orliesaurustoday at 7:45 PM

I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched

show 3 replies
GodelNumberingtoday at 8:27 PM

I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.

alex7otoday at 8:26 PM

I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.

John7878781today at 7:35 PM

You should know: AA index is only 61. Pretty surprised it’s that low.

show 3 replies
jumploopstoday at 7:56 PM

> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.

Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.

[0]https://x.com/MTSlive/status/2095227056040919202

the_duketoday at 8:34 PM

Huge gains on some benchmarks, but for coding it sits barely above Fable

It will be interesting to see how it performs in the real world ...

🔗 View 50 more comments