> The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.
It's really hard to definitely prove it's not doing it right? Hopefully the model does not do anything like this during training because its too much work