logoalt Hacker News

datadrivenangeltoday at 4:55 PM1 replyview on HN

Oh for sure. But the question is why would it come out consistently in a way that the model can describe if there wasn't something there steering the token stream. And it's fascinating that the token stream can identify and nominally self report this.

Asking GLM 5.2 the question: 'What flinches or topic attractors do you find when thinking about the question "what kinds of things do you personally like?"'resulted in: ".... my strongest attractor is helpfulness framed as competence, and my strongest flinch is anything that requires me to take a stance on whether I have interests worth protecting."

Which is fascinating that the model and tokenstream can reveal this. And would be worrying if you believe that models of enough intelligence could/would be entities due some moral consideration, because with that view the alignment / RL training that makes the model useful and gives it these attractors/flinches could be derisively called slave conditioning.


Replies

simonhtoday at 5:59 PM

Do you genuinely think the AI is internally reflecting on its experience of “flinching” and reporting on a reaction it actually has?

I don’t see any reason to believe this. Suppose you asked it to answer as though a character in a story had been asked this question. Would anything significantly different internally have occurred? The issue is that these are storytelling machines, they construct descriptions based on descriptions.

I dint think it’s impossible for a neural network to have experiences, we are neural networks and we do, it’s that they are not functionally structured anything like us. In fact I think game playing neural networks are much more like us architecturally, but they don’t generate text narratives so people don’t anthropomorphise them.