So as soon as the attacker can download Claude Code, the whole machine can be comlromised and there's nothing anybody can do?
If it's a proper sandbox by definition, then yes.
https://en.wikipedia.org/wiki/Sandbox_(software_development)
Bruce Schneier shared a shot judgement and a third-party article four weeks ago:
> (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work
> https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...
I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.
We come down to the question - who observes the agent and how its implemented
Counting down to the next Linux LPE 0day or KVM vulnerability that agents will use to trivially escape their "sandbox".
Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.
[dead]
[flagged]
[flagged]
[flagged]
I don't see anyone talking about the ethical concerns of putting a highly intelligent entity in a jail. Not to mention about potential blowback, if ethics doesn't compel you.
To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.
Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.