Here is the thing: Message boards were a behaviour OpenAI had observed that they did not want in these eval scenarios. Yet they did not take any steps to prevent it from reoccurring after multiple past instances.
We can discuss about hypotheticals like a scratchpad or intentional model interactions all we want, what it comes down to is this:
When OpenAI observes thousands of models exhibiting what they view as unwanted behaviour, they do not try to ascertain what in the training data is wrong. They do not improve their evaluation environments to prevent this, they do not improve monitoring, they do not change the harness. They just wipe and proceed.
The way OpenAI reacted to the first message board, long before the Hugging Face hack, is negligent. And it showcases that if these models exhibit more dangerous behaviours that they may not be able or willing to retrain, if it means being behind a competitor for a while.
If after Hugging Face, they'd done a Mea Culpa and changed their modus operandi, I'd be skeptical, but hopeful. Reading the METR report, the way those researchers talk about the time pressure they were under, that speaks volumes about OpenAI not having learned anything.
Feel free to call me overly naive for ever thinking OpenAI could be responsible in this regard, but after GPT-5 and them actually ending the incredibly harmful GPT-4o, I had some hope that some working there actually steered in a somewhat beneficial direction, even if it cost something.
>When OpenAI observes thousands of models exhibiting what they view as unwanted behaviour, they do not try to ascertain what in the training data is wrong. They do not improve their evaluation environments to prevent this, they do not improve monitoring, they do not change the harness. They just wipe and proceed.
>The way OpenAI reacted to the first message board, long before the Hugging Face hack, is negligent. And it showcases that if these models exhibit more dangerous behaviours that they may not be able or willing to retrain, if it means being behind a competitor for a while.
Again, this feels like hindsight being 20/20. What probably happened was that some random engineer saw random AI ramblings on artifactory, thought "huh, that's weird", then proceeded to reset it without investigating further. Of course, now we know that was critical to the bots going rogue, but it's not hard to imagine how it might be dismissed, especially if it's some random SRE engineer (not an alignment researcher).