logoalt Hacker News

zaptremyesterday at 11:14 PM1 replyview on HN

Can you explain how the above event doesn't count as evidence alignment is an actual risk?


Replies

JumpCrisscrossyesterday at 11:28 PM

> Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)

show 3 replies