So the models will not only be using more and more Neuralese in their CoT (like GPT-6), but different agents will also be able to communicate with each other in Neuralese. It's not looking good for monitorability.
Is Neuralese in no way decodable into a human-interpretable system? Genuine question -- I don't know the answer.
This is from like a year ago.
Is Neuralese in no way decodable into a human-interpretable system? Genuine question -- I don't know the answer.