The models "go rogue" because they are not sufficiently sandboxed. Arguably there is no criminal intent on the side of OpenAI in all of those cases. And in at least one case the agents were operated by other companies.
So it looks to me that any liability would be civil in nature and given the actual damage done pretty limited.
Why would you need to sandbox them? These models are apparently trained to do this, how about we just don't include that training data?
Sandboxing is just an endless race to patch holes and you can only sandbox the agents so much before they become useless. Unless you screen the training data and avoid teaching the LLM about "hacking" and looking for API keys on Github, you'd have to completely disconnect your agents from the internet and file system. At that point agents starts to be rather useless. All the talk about sandboxing and guardrails is just corporate/management speak for we don't want to fix the core problems in our product.
In the US, isn't hacking and avoiding security restrictions online going to be wire fraud, regardless of your intentions and actual damage? That's not a civil matter. What you could do in that case is to go after the user operating the agents. That would make the user act as the emergency break for otherwise uncontrollable agents.
The first time it happens, “there is no criminal intent” may carry some weight. After tens of thousands of instances, a lot less so…
Isn't there the concept of criminal negligence?
Can't they just use this powerful AI to build a sufficient sandbox?
Seems like the first thing one would do.
Let’s start with an investigation of the company and see if that’s the case or not?
So you know their internal thoughts? What they did? Criminal intent does not matter. These are the most knowledgeable people on earth, supposedly, yet, they are beyond negligent?
Which is it?