>If the recurrence step is no longer in English, we can no longer monitor intent, and can only observe whether a model is safe through behavior.
So, you mean, like another human person?
> So, you mean, like another human person?
No human is vastly better than all humans at all cognitive tasks.
Humans can't think 100 times faster than humans.
Humans when interacting with computer networks have limitations on how fast they can do so.
Humans have millions of years of evolution, and thousands of years of cultural evolution, in creating ways of detecting and alleviating dishonesty and non-alignment with other humans; much of this will not work with AIs.
Most people do not hide intent, we also write things down at work, eg. a ticket in kanban
We could never rely on scratchpad to reflect true model deliberatons.
I don't know about you but I don't tend to let other humans have access to my computer directly to do stuff for me