logoalt Hacker News

GPT-6 Astra

1036 pointsby kibaetoday at 6:41 PM751 commentsview on HN

System Card: https://deploymentsafety.openai.com/gpt-6-astra

Related ongoing threads:

OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691

GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147


Comments

dgellowtoday at 7:38 PM

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

Wait, what? Am I understanding that correctly? That sounds really bad

show 4 replies
jumploopstoday at 7:56 PM

> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.

Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.

[0]https://x.com/MTSlive/status/2095227056040919202

Robdel12today at 8:35 PM

I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.

So, folks that have actually used this already, what’s it actually like?

GodelNumberingtoday at 8:27 PM

I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.

alex7otoday at 8:26 PM

I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.

itissidtoday at 9:34 PM

All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.

Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.

dangtoday at 8:24 PM

Argh! I hit a wrong keyboard shortcut and moved the entire thread.

Please stand by... it will all come back shortly

show 2 replies
KolmogorovComptoday at 7:50 PM

GPT-7 Zeneca

show 2 replies
jerrygensertoday at 6:44 PM

> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.

show 1 reply
the_duketoday at 8:34 PM

Huge gains on some benchmarks, but for coding it sits barely above Fable

It will be interesting to see how it performs in the real world ...

BrokenCogstoday at 8:39 PM

GPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw

simonjgreentoday at 7:44 PM

https://youtu.be/1QNsdr-Qx_I?si=coXwStCl7clpGVC1 Launch video

show 3 replies
smashers1114today at 8:10 PM

I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.

wiseowisetoday at 9:29 PM

Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?

show 1 reply
KronisLVtoday at 9:15 PM

It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.

gizmodo59today at 7:47 PM

99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.

cromkatoday at 9:40 PM

Surprised they haven't reset Codex usage on this occasion.

show 1 reply
sashank_1509today at 8:33 PM

Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see

carlos-menezestoday at 9:45 PM

The Kart Racer game is easily breakable if you spam the spacebar.

AGI!

tekacstoday at 7:47 PM

https://developers.openai.com/api/docs/guides/latest-model

The docs page has a bunch more interesting details, including for example async tool calling!

brindidriptoday at 9:09 PM

Cool, I don't really care anymore.

hannofcarttoday at 8:19 PM

What does 'Astra' here mean? Surely they must be referring to the Latin word.

Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.

show 3 replies
mvkeltoday at 8:32 PM

The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.

If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.

Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.

hazelnuttoday at 8:20 PM

Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.

alpinemantoday at 8:56 PM

That Astra ‘city scene’ is about as creative as Doha in real life (not very)

kegs_today at 7:37 PM

I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come

show 11 replies
sharmajaitoday at 8:46 PM

Really feels like AGIPO is here.

udbhavstoday at 8:22 PM

Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.

foundOpenRighttoday at 8:16 PM

1:15.425 on Sunset Cove beat my record

ianm218today at 8:26 PM

I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.

gekoxyztoday at 7:54 PM

HTTP 500 for me on the announcement page :(

show 1 reply
Oblunesstoday at 9:41 PM

That seems promising ?

E-Reverancetoday at 8:30 PM

At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete

Rover222today at 9:16 PM

Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.

I hop models at will, and have done 90% of my work on OpenAI models since sol came out.

prometheus1992today at 7:54 PM

this is crazy! can't wait for the 27B distilled version of this.

HardCodedBiastoday at 9:51 PM

Even though the model is clearly wonderful the launch video is an abomination.

That gives me hope that there is still areas to improve.

What a bad launch video. Hilarious.

What a powerful model.

HSOtoday at 10:03 PM

people are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla)

the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing

agi deus ex machina descending from the icloud ftw!!!

pathetic :)))

rbrevetoday at 8:55 PM

Where is the cure for cancer?

show 2 replies
semiquavertoday at 8:09 PM

Guessing this one will never show up in cursor…

alex7otoday at 8:31 PM

Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad

brcmthrowawaytoday at 8:42 PM

Anthropic in tears today.

firemelttoday at 7:56 PM

damn seems I should hold off my claude subs

wahnfriedentoday at 6:45 PM

They're just announcing later availability. No launch.

show 2 replies
jiraiyasarutobitoday at 8:17 PM

It saturated most benchmarks. WTH

saaaaaamtoday at 7:37 PM

Pelicans please

show 1 reply

🔗 View 41 more comments