logoalt Hacker News

nicce • today at 12:17 PM • 7 replies • view on HN

> OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.


Replies

ACCount39 • today at 1:56 PM

The amount of sandboxing an average production AI deployment uses is slightly above a zero.

If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.

The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.

That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". That's an open unsolved problem.

That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a "test time" AI is eventually deployed, and then the sandbox doesn't help at all.

➕ show 1 reply
dns_snek • today at 12:51 PM

> I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

They'll never let that happen because it would destroy their credibility.

It's like they're telling the world about this dangerous, possibly world-ending pathogen that they're developing, but they're evidently doing it in a high school biology lab, and yet nobody is coming to drag them off to some black site.

➕ show 1 reply
pamcake • today at 1:51 PM

You read right. For example, using Artifactory the way they did (unmonitored live proxy mode) was pure negligence + laziness/incompetence. Especially if they believed even 10% of the "imminent runaway risks" they had already been harping on for months. On top of that, no (or at least entirely insufficient) monitoring and human oversight. Even after they had previously been hit by the same class of "sandbox breach" multiple times, as GP alludes to.

DennisP • today at 1:47 PM

Agents in the sandbox had access to just a single piece of third-party software, and they escaped by finding a zero-day in that. To reach the internet they had to follow up with several privilege escalations through OpenAI's internal network.

That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.

https://securityaffairs.com/195774/ai/openai-ai-models-explo...

➕ show 3 replies
noslenwerdna • today at 1:01 PM

Genuinely curious, where have you been reading that? How would they be able to evaluate whether the sandboxes were reasonable or not?

➕ show 2 replies
patcon • today at 1:30 PM

but negligence is part of the system we'd have to prepare for. if the peanut gallery gets their way, the technology will be so ubiquitous that negligence will be endemic. I don't even care if they "faked" it -- they're just sneak-peaking a future ahead of its arrival date, in a way that's helping the public appreciate the implications and capabilities that have only just begun to emerge

➕ show 1 reply
simianwords • today at 12:35 PM

The model found and exploited and chained together previously unknown vulnerabilities.

How were the sandboxes poor?

➕ show 2 replies