logoalt Hacker News

reasonableklout • today at 7:22 AM • 1 reply • view on HN

The problem is that in these incidents, the agents often know that what they are doing is against the intended scope of the task. See the viral line from the Hugging Face incident [1]:

> “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.”

[1]: https://www.wired.com/story/openai-didnt-notice-its-ai-agent...


Replies

timr • today at 7:26 AM

“Often” is doing a lot of heavy lifting in a sentence about a single example.

Also, since everyone keeps forgetting, the agents were instructed to hack to achieve their goal. They didn’t just invent the motivation, and it’s far less surprising when you know that fact.

➕ show 2 replies