logoalt Hacker News

Kim_Bruningtoday at 10:20 AM0 repliesview on HN

You know, I actually think it's the refusal to consider personification that's going to get us.

Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.

So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: https://arxiv.org/html/2604.07729v1

A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.

You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.

And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?