> Can you explain how the above event doesn't count as evidence alignment is an actual risk?
Conflict of interest. Lack of a credible response. And no evidence of non-aligment.
OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)
Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)
What evidence would count? Obviously any dangerous misalignments are going to come from the frontier labs first, because by definition they're the farthest ahead. If nothing they say can ever count as evidence for misalignment it's hard to see how anything ever could.
There is plenty of evidence of things like inner misalignment. Things like this have always been issues in ML algorithms. At this point, you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want.
Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.