> In adversarial settings (where we push the model to evade our monitors)
...why exactly are they training for that?
Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
presumably that's a safety evaluation not a training setting