logoalt Hacker News

ben_w • today at 8:01 AM • 1 reply • view on HN

A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.

"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".

(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).


Replies

dns_snek • today at 8:58 AM

> already understood (we can tell because they wrote it down)

No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?

➕ show 1 reply