logoalt Hacker News

gruezyesterday at 5:08 PM3 repliesview on HN

>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

>3) Huggingface has no security and the AI broke in using standard script kiddie methods.

Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?


Replies

orbital-decayyesterday at 7:36 PM

No, a hacking benchmark was exactly what it was tasked with. It wasn't its way to bake a cake.

arjieyesterday at 7:40 PM

The home directory rm situation also adds credence to this take. The Claude series is much better aligned in comparison.

wonnageyesterday at 5:15 PM

Didn’t they explicitly remove alignment guardrails for this test? From the press release:

> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

show 2 replies