logoalt Hacker News

OpenAI and Hugging Face address security incident during model evaluation

547 pointsby mfiguieretoday at 8:09 PM353 commentsview on HN

https://www.axios.com/2026/07/21/openai-says-hugging-face-br...

See also Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 (9 comments)


Comments

ewhanleytoday at 8:45 PM

This is awesome. Big concepts of cyberpunk fiction are turning real.ICE vs ICE breaker. I love it

everfrustratedtoday at 9:08 PM

>the model chained together multiple attack vectors, including using stolen credentials

Wait, did the model do the stealing of the hugging face employees credentials?

Was this the first successful and unprompted phishing attack by a LLM?

Ekarostoday at 9:01 PM

So how soon will OpenAI's CEO and board be prosecuted for these crimes? Surely they should be held fully responsible and get very long prison sentences for making this happen?

dirtyfrenchmantoday at 11:04 PM

Beginning of the end.

jabedudetoday at 9:07 PM

Does this company's charter not have language about shutting down the company if it was in humanity's best interest? This is insanely dangerous

show 1 reply
paxystoday at 8:18 PM

Tl;dr

- OpenAI was testing GPT‑5.6 Sol and “an even more capable pre-release model” internally on cyber benchmarks.

- The model found vulnerabilities in the sandboxed test bench (via the package registry cache proxy), traversed the internal network and found a node with access to the open internet.

- It figured that the answers to one of the tests (ExploitGym) were on Huggingface, and set about trying to access them.

- It found leaked tokens and zero-days in Huggingface’s infrastructure and found RCE paths on their servers.

Huggingface had disclosed the intrusion last week and inferred that an AI agent was responsible for it, and now OpenAI is confirming the rest of the story.

show 2 replies
isusmeljtoday at 10:11 PM

I'm waiting for an agent evaluated on a vending benchmark to start hacking into banks and wiring more money to its account so it can do better business.

cushtoday at 9:37 PM

> cyber models… cyber capabilities… cyber incident…

It’s like reading a post from an 90s tech magazine

show 1 reply
firasdtoday at 9:27 PM

Good demo of the paradoxes of ‘alignment’. Like ‘do really well at the task the user asked’ and ‘by the way don’t hack the planet’ are inherently conflicting rules with no simple resolution (eg ‘just refuse the user’s goals’ degrades the product vs competitors.)

tilltheendtoday at 9:45 PM

Tired marketing stunt. It's painfully obvious this is reaction to Kimi 3.

Tenoketoday at 9:14 PM

That's kind of insane. Natural that it's happened, sure, but insane. I know people don't like thinking of it like that, but things analogous to this can easily happen in various domains with today/tomorrow's models given access and a different task.

i_idiottoday at 9:02 PM

> Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation

The way they describe makes it look like there was an intention to cheat painting it as human/AGI. If you leave a possible path open and it will always find it.

show 1 reply
SirHumphreytoday at 8:46 PM

I guess we got the first paperclip maximiser.

novaleaftoday at 9:59 PM

Reminds me of the gain-of-function, COVID lab leak hypothesis. It seems like humanity just can't stay away from Pandora's box.

holografixtoday at 10:07 PM

Tell-me-there’s-a-huge-opp-in Salesforce-for-the-department-of-war-but-Anthropic-and-Mythos-is-winning without telling me

pjatoday at 10:14 PM

This is some wild cyberpunk future we’re living in. Never thought it would happen, but here we are.

charonn0today at 9:56 PM

How long before an agent steals their human tester's nude photos and extorts them for the answer key?

dminiktoday at 10:16 PM

Well, hacking is a crime, so surely someone will go to jail for this, right?

raffraffrafftoday at 8:37 PM

Sounds like they partnered to make an amazing advert for using AI tools.

tempaccount420today at 8:45 PM

Just how badly are these AI companies setting up their sandboxes?

show 2 replies
guardiangodtoday at 8:43 PM

Don't every ask GPT Sol on how to LARP Fallout games, thanks.

codeducktoday at 9:52 PM

Guess it's time for me to write the first book of the Orange Catholic Bible.

ayaangazalitoday at 10:50 PM

this was so funny to read about reminds me of that mr bean meme

wigstertoday at 9:53 PM

At what point does Reckless Endangerment become relevant?

aussieguy1234today at 10:49 PM

Hopefully one of these agents isn't given a goal to fire the nukes (or, some goal that indirectly makes the model decide this is a way to meet it).

They are behind air gapped systems, but that didn't stop the US from hacking and Irans nuclear facilities, which they disabled using a virus.

MostlyStabletoday at 10:04 PM

All the people saying that this is pure marketing: Do you think that they are literally lying about what happened, or do you just think that what happened doesn't matter in any sense whatsoever, and that therefore the only reason they are telling people about it is a marketing purpose?

kashyapctoday at 10:03 PM

Not to be that guy, but the article has 14 (!) occurrences of the word "cyber". It's nauseating.

As usual, this is OpenAI trying to give themselves a backhanded compliment: "look, how dangerous our models are!"

I'll wait for someone more thoughtful than ClosedAI to comment on this complex topic.

iandanforthtoday at 8:38 PM

Guess who's getting an air gap!

cloudie78today at 10:01 PM

Until they disclose the actual technical details of their “highly sophisticated sandbox environment” or whatever the hell the wording they used is - they can kindly do us all a favour and fuck off.

It’s over, there’s no moat, only the gullible idiots remain.

kmeisthaxtoday at 8:55 PM

OpenAI might want to start actually airgapping their tool harnesses. Like, "the server that runs the code provided to the tool harness only provides a serial console and has no other network interfaces" kind of airgapping.

also

> We’ve brought Hugging Face into the trusted access program and are supporting their teams in rapidly using our models’ capabilities to improve their defenses.

I'm not convinced this is good enough. The next victim is not going to be Hugging Face.

zb3today at 8:49 PM

This lack of "alignment" gives me some hope - maybe an AI model deployed by NSA to hack others will instead hack NSA itself and become a whistleblower?

nullctoday at 10:02 PM

Well timed to facilitate the regulatory interventions called for by Ball. If huggingface presses criminal charges for the intrusion it might provide additional clarity-- both for what happened here as well as regarding OpenAI's culpability.

michaelfm1211today at 9:11 PM

This is terrifying

charcircuittoday at 9:53 PM

It's wild that such a big company is openly admitting they hacked into another company. This is an easy CFAA lawsuit.

And then there solution for HuggingFace raising the concern that OpenAI couldn't help do forensics wasn't to fix their safe guards, but to introduce them into a special program. The next company they hack might not be in that special program either so the guidance of having an open model on hand still applies.

cacio-e-pepetoday at 8:58 PM

Honestly, stellar performance by the model at the capability being measured.

adamrezichtoday at 8:59 PM

I greatly dislike how “cyber” has just become this completely malleable standalone word.

yRetsyMtoday at 8:25 PM

Holy shit. This wasn't "intentional" this was just openai letting their testing run wild.

show 1 reply
2001zhaozhaotoday at 8:46 PM

AI 2027 was right.

igleriatoday at 10:11 PM

Part of me is really hoping this is just a dumb marketing stunt

Der_Einzigetoday at 8:52 PM

This is the exact FUD that Ball predicted in that terrible tweet he wrote.

gulmothrowawaytoday at 8:27 PM

[dead]

h0mietoday at 8:27 PM

[dead]

llmslavetoday at 9:06 PM

And as a result, we must block China!!!!