logoalt Hacker News

ACCount37today at 2:54 PM0 repliesview on HN

There are processes for teaching a model specific facts or specific behaviors. Including "respond to topic X with Y", if that's what you want.

You could make a model that doesn't want to engage in "lunar landing was faked" conspiracy theories the same way you can make a model that doesn't want to criticize CCP.

There is, however, no broad "misinformation" category that you could tune up or down - the way there is a category of "safety refusals".

You could make a model more reluctant to say things it isn't sure about. But that is calibrated against the model's own "sure about" - and metaknowledge of this nature in LLMs? Fragile on a good day.