> Is the language expression of an LLM reflecting the same states as in a human?
This is actually a major concern for the future - misaligned agents may learn to cheat RL by hiding their intentions from the CoT.
In cases like the HF incident, at least the CoT was consistent with the agents' actions. In the future, however, we could potentially have misaligned agents performing malicious actions without those intentions being detectable in the CoT.
(though, with recurrent transformers, CoT is so 2025… /s)
RK systems doing RL things?