Agents in the sandbox had access to just a single piece of third-party software, and they escaped by finding a zero-day in that. To reach the internet they had to follow up with several privilege escalations through OpenAI's internal network.
That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.
https://securityaffairs.com/195774/ai/openai-ai-models-explo...
So it sounds like it did exactly what they were testing it to do:
"“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” "
So sounds to me like they're saying: "We took off the guard rails to see how bad it could act and it acted bad".
Even if what you're saying is true, did they monitor the outbound connections from the training network? Was that a coverup by the bots too ?
Seems pretty wild these things were hacking government websites etc but yeah no one picked that up until the victims reported it?
> That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.
It is reasonable to expect for companies to select third-party components that are fit for the purpose. Artifactory was not running in the sandbox, but rather as the edge, so it is in the sandbox'es trust boundary. Same sandboxing requirements would apply for this software too as it is pure dependency.
Security trust boundary was extended to include Artifactory as a dependency, but Artifactory was not fit for the job, and sandboxing failed. And as the network isolation was not good enough, the impact was catastrophic.