logoalt Hacker News

gck1yesterday at 11:21 PM8 repliesview on HN

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment

> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations

> we identified three incidents

> The incidents involved three different Claude models: [...] and an internal research test model

This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.

I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.


Replies

simonwyesterday at 11:25 PM

I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations!

The hacks weren't particularly impressive either:

> [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities [...]

show 7 replies
strictneintoday at 12:08 AM

I'm cynical as well, but the logical thing for them to do after the OpenAI/HF incident was to look at their systems for similar activity.

If they hadn't published this and instead it leaked out in two months we'd be slamming them for that as well.

They're stuck between a rock and a hard place, although they kind of put the rock there.

nissa-seruyesterday at 11:37 PM

No - the pain of the person writing that post comes through in the words; shipped quick, lots of stakeholders, single owner i bet, "how the fuck am i supposed to toe all these lines simultaneously"

_dain_today at 12:39 AM

Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme?

This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident.

GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO.

---

Put another way, how would you have done the write-up about one of these breakout incidents, if you were in an Anthropic/OpenAI employee's shoes, and (by hypothesis) your intent were not "marketing"? And in a way that doesn't trigger the "it's all marketing" HN top-ranking comment?

show 4 replies
protocolturetoday at 3:00 AM

>This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.

This was my immediate thought.

skeptic_aitoday at 12:11 AM

For the big safety guys to only investigate this either means are incompetent or malevolent. Which one?

Tip: the people working there are the top 0.001% smartest in the world

show 1 reply
patconyesterday at 11:29 PM

[dead]

sscaryterryyesterday at 11:56 PM

Just trying to have the limelight back on them. Utter and complete bullshit. Just like the OpenAI "incident".

A human instructed an LLM to perform a certain task, I'm sure (unless I've really lost my mind) these follow instructions, with some judgment, in a loop.

Given all the other negative publicity around industrial espionage, with at least OpenAI being fingered, it would not surprise me if this was intentional.

(Edit): In case it wasn't clear. I fully agree with the op.