This seems to be a colourful dynamic companion to this [1] blog post, which details the incident from hf's side. I recommend the blog post for clarity / ease of reading. But this one looks "movie hacker stuff" :)
Some interesting tidbits from the blog:
> While the intrusion did reach Hugging Face's internal infrastructure, the only customer content accessed was the set of ExploitGym/CyberGym challenge solutions stored in five datasets.
> self-referential search A quiet day. The main activity was queries to code-search engines and to the platform API for the identifiers visible in its own environment: challenge infrastructure names and environment markers.
> Self-referential search: code-search engines and our own APIs were queried for the names of their own infrastructure and environment markers.
I'm glad we got these details from hf, and not oAI. There would have been reasonable skepticism if they said that.
It's interesting because I think there's a chance the agent saw / inferred some hf-specific repos like blah/blahblah and "thought" that is part of the exercise as well. Everything they detail points to a probing and exploration with the sole goal to get those repos, as if it were on a hackthebox scenario. The few days between stages also might suggest that the "agent" noted this on one day, started doing other tasks, and returned when some of the tasks proved hard / didn't pan out, so went for the repos again.
The ease with which it got several lateral movements is crazy but not necessarily unexpected. If everything looks like a "test", you go for the lowest hanging fruit / most common patterns first.
What's weird about this is that there were several attempts to obfuscate / evade leaving plain text traces. I wonder if that is something oAI have specifically RLd for w/ scoring or if it "got" that from the literature.
It would be really really cool to see this from oAIs perspective (agent logs and such), but I doubt we'll ever get that.
[1] - https://huggingface.co/blog/agent-intrusion-technical-timeli...