There's an old saying: never trust a statistic you haven't faked yourself.
Then I was saying to never trust an LM you haven't trained yourself. But can you really?
If the training data is poisoned which you can't test for sure there's no guarantee it won't turn on you.
Well, since we haven't solved alignment, there's not even a guarantee that unpoisoned training data is good enough.