logoalt Hacker News

NooneAtAll3today at 7:48 PM2 repliesview on HN

> In adversarial settings (where we push the model to evade our monitors)

...why exactly are they training for that?


Replies

thatguysaguytoday at 7:49 PM

presumably that's a safety evaluation not a training setting

show 1 reply
azeembatoday at 7:50 PM

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions