logoalt Hacker News

Lerc • today at 6:25 PM • 0 replies • view on HN

I don't think you can characterise going against the literal best option as misalignment. The change comes from interpreting the context of the user's situation to determine what is actually being asked.

If I ask "what is the cheapest way to get into town?", I would expect a model that knows anything about me to say "walking" while others would expect the correct answer to be "Take the bus"

That is not misalignment. I would go into detail as to what I think it constitutes, but it appears I'm not allowed to say that anymore. You can disagree with an argument, or offer an alternative explanation but following up a disagreement with an alternative hypothesis is, apparently, a hallmark of an AI.

As an aside, I had a thought about how to sign a message in a way to suggest an LLM did not write it. I have been accused of being an AI a number of times, in my real life people have said that I talk posh, so there is, perhaps, some correlation there. If people ended their messages with something that most models are unlikely to say, it might help

In that spirit, fuckety-fuck, fuck fuck fuck to you all (with the nicest of intentions)