I'm not skeptical that this attack happened, I'm skeptical that the model's prompt was truly just "solve this benchmark" and nothing more.
I'm also trying to figure out why OpenAI put out a press release about this. In what way is this not admitting to a federal crime?
Because this is amazing PR? Just following the Anthropic rulebook.
yup I am dead certain that they realized the whole 'our AI is too dangerous' punchline is too played out and they needed something that makes actual splash. Also this incident would serve as a foundation to ban open-weight models because 'with great power comes great responsibility' or some other BS like that. because after all, unwashed masses cannot be expected to be 'responsible' with top of the line intelligence.
I expect these type of hacks to continue till they IPO. after that real public company liability will start taking over.
If Hugging Face and the feds are one step away from discovering your attack what other option do you have but to come clean?