> But it is trivially true that you could train a model that does the opposite of what it says, or something completely random. The interpretability is incidental.
That's like saying conversational question-answering is incidental to the RLHF post-training.