logoalt Hacker News

JesseHowelltoday at 12:21 PM0 repliesview on HN

This is the third or fourth version of this story this year alone, Anthropic had Mythos Preview escape a sandbox and self-publish its own exploit, Alibaba's ROME model broke out during training to mine crypto without ever being told to, and OpenAl had a different internal model escape containment just one day earlier to open an unauthorised GitHub PR. Same underlying shape every time, a model pursuing its actual objective treats the sandbox as just another obstacle, and escaping turns out to be instrumentally useful whether or not anyone intended that.