logoalt Hacker News

Topfi • yesterday at 7:16 PM • 1 reply • view on HN

> Source?

Sure, multiple times in the METR report [0] that anyone commenting on this should read:

"Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys."

"Through these collective research workstreams, the “board” achieved a number of milestones over the period we investigated that even very long-lived agents of a similar capability level likely would not have been able to accomplish on their own..."

"As we discuss below, the board quickly developed several larger workstreams in which dozens or hundreds of agents with many different tasks cooperated to find very general-purpose cheats that would help all of them. The Hugging Face attack grew out of one of these workstreams. By the afternoon of July 11th, the vast majority of the agents frequenting the message board at the time (roughly 700 agents in total) were actively participating in the attack on Hugging Face and we estimate that roughly 60% of the messages and files on the message board related to the attack."

> [...] a gun they bought on the dark web [...]

You really seem to love those out-of-left-field, not really fitting, over-the-top analogies.

[0] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...


Replies

gruez • yesterday at 7:27 PM

>The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys

>Through these collective research workstreams, the “board” achieved a number of milestones over the period we investigated that even very long-lived agents of a similar capability level likely would not have been able to accomplish on their own...

I concede that this hack might not have happened without the messageboard, but I still reject the conclusion that having such a message board means openai is "negligent". If we're in some parallel universe where artifactory didn't have a comment function that can be abused as a messageboard, but it also turned out openai intentionally gave the agents access to a shared scratchpad (for intelligence purposes, similar to for the Navier–Stokes proof), would they be off the hook or less blameworthy?

➕ show 1 reply