>Generally, it is hard to imagine how neuralese should work given that models are pre-trained on naturalistic documents:
If you look at any paper/blog etc detailing Reasoning RL runs, they'll tell you the same thing. 'Thinking' text trends towards unreadable gibberish (for humans) unless you reward for it. Even then, take a look at the scripts in the Huggingface incident and most of it is dense stuff that's hard to parse. They had to rely on agents to make sense of it.