logoalt Hacker News

jubilantitoday at 7:53 PM0 repliesview on HN

> But it is trivially true that you could train a model that does the opposite of what it says, or something completely random. The interpretability is incidental.

That's like saying conversational question-answering is incidental to the RLHF post-training.