What exactly prevents anyone worried about this to build an LLM that can decode the neuralese to english and use it to monitor what the model is doing?
Or is the neuralese some sort of irreversibly encrypted data set that only an LLM can "understand" and that can never be translated back to English?
Turtles all the way down:
https://www.astralcodexten.com/p/elk-and-the-problem-of-trut...
How do you trust that LLM?
It's very hard in practice. Even without optimization against it, it's about as hard as understanding an activation layer today and will probably get harder in the future.
There's also nothing stopping the secondary LLM from confabulating bullshit and there are less checks on it than CoT (harder for either humans or other models to externally verify).