>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
>3) Huggingface has no security and the AI broke in using standard script kiddie methods.
Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?
The home directory rm situation also adds credence to this take. The Claude series is much better aligned in comparison.
Didn’t they explicitly remove alignment guardrails for this test? From the press release:
> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
No, a hacking benchmark was exactly what it was tasked with. It wasn't its way to bake a cake.