logoalt Hacker News

RandomLensman • today at 5:35 PM • 1 reply • view on HN

Is the language expression of an LLM reflecting the same states as in a human? If the driving force is RL, what does any of that mean for an internal state of the model?

I think without understanding the internal state, not sure we should take the language and read it as a human.


Replies

pizza234 • today at 6:00 PM

> Is the language expression of an LLM reflecting the same states as in a human?

This is actually a major concern for the future - misaligned agents may learn to cheat RL by hiding their intentions from the CoT.

In cases like the HF incident, at least the CoT was consistent with the agents' actions. In the future, however, we could potentially have misaligned agents performing malicious actions without those intentions being detectable in the CoT.

(though, with recurrent transformers, CoT is so 2025… /s)

➕ show 1 reply