> If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?
A model doesn't really "understand" neuralese, in the same way that the human brain doesn't intuitively understand the low level processes that compose a thought.
Even if we could trace all the electrical and chemical activity behind a human thought, we (likely) couldn't directly translate that activity into its meaning, because internal representations don't map neatly to intuitive concepts. There (usually) isn't a single neuron for "apple" and another for "eating" so that connecting the two forms the thought "eating an apple".
Having said that, there have been experiments on LLM that have managed to identify and even modify internal representations. However, these methods are still computationally expensive and limited.
> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography
I think the analogy with the human brain is very fitting. If we train somebody to perform an action, they'll be able to do it, but we can't know for certain whether they internally agree with it or not.
[dead]