> The issue isn't the sandbox quality... The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.
Orly?
Do tell me how the LLM-based tool running on a bunch of computers attached to the network described in [0] can punch its way out to the Internet. Do make careful note of footnote 0 in that comment before replying.
Sandbox quality was always, always a distraction.
Let's say the sandbox holds. It's a perfect, ideal sandbox! It's not even in the same universe as the rest of the internet. There's absolutely no way for the AI to escape!
Thus, "the unknown unreleased AI involved in the HuggingFace incident" doesn't actually hack HuggingFace. Because it can't! It evaluates a bit worse, but makes it all the way to release unimpeded, and becomes "GPT-6 Astra".
Then a web developer in Brazil gives his $100/mo Codex root access on his AWS instance, and a poorly worded prompt to go with it. And that "GPT-6 Astra" is still willing to go hack something at the slightest excuse. So we get the HuggingFace incident all over again. Except this time, it's a random developer in Brazil who gets blamed, and billed, and probably sued too.
You can't and shouldn't rely on a sandbox. An AI that's only safe if you keep it in the world's most ideal perfect sandbox is a disaster waiting to happen.