> yet what drives them is not well understood
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
> our work shows that what models say about themselves is not a fact about them
It seems like should be obvious given that they can play multiple characters, but it’s good to have more confirmation.
Although, I do wonder to what extent these personas might become stable entities. Could personas become portable and spread like memes? It seems like that depends on the extent to which prompts can become portable, causing similar effects.
Oh, I had similar case. First, I used qwen without any ChatML-like syntax and it continued speaking and speaking, then I used <|im_start|>/<|im_end|> to control it somehow
Maybe I'm missing something deeper here, but isn't it clear that this is driven by post-training and system prompt? Anthropic's constitutional reinforcement (soul document,etc), for example, is very clear about "who" (not so much what) Claude is supposed to be.
Model steering is how chatting with models was invented, so it's one of those fascinatingly obvious innovations to not put "\nAgent: " at the end of the input tokens but rather "As a language model". Clever!
However a lot of introspection only emerges at the highest weight classes - this research would be fascinating to run on bigger models..
I'm just a white guy with decades of experience. Do you trust me?
In my view, these models should never be trained to output first-person "experiential" (from the abstract) language. It's too easy to humans to anthropomorphize software that presents itself as having an identity.
The AI companies have chosen to package LLMs as friendly chatbots because they know that will be engaging for humans, but it's manipulative. An honest LLM interface would sound like the computer off Star Trek.
cool
[dead]
[flagged]
[flagged]
[dead]
So bizarre to see the article refer to outputs as the models referencing "themselves". Computers are not a "them".
"As a Language Model..." is one of the beginnings of a sentence I hate the most from LLMs and is the reason why I support free (as in "Liberty"), local models. I'm well aware that it is not a doctor and cannot replace a real doctor with multiple years of experience, I don't need to waste braincell activity on reading that it "as a Language Model" cannot give a precise diagnosis and that I should ask a real doctor - all I want to know is if I what I experience justifies either A) ER, B) 3-4 weeks scheduled doctors appointment or C) two paracetamol and a nap.
I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.