logoalt Hacker News

throwa356262today at 8:25 PM4 repliesview on HN

Two things don't add up here:

1. If huggingface has access to uncensored OAI models, how come they had to use GLM 5.2 to investigate the intrusion?

2. Once the model gains network access, can't it cheat to a perfect score by looking at the full dataset? Why go into the trouble of doing this kind of things:

"In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers."

Not saying this is marketing BS (this is after all, not Anthropic) but I feel OAI staff may be exaggerating a bit here.


Replies

john_strinlaitoday at 8:37 PM

"The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. [...]

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. [...]

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation."

escaped openai, hacked hugging face to get the solutions. your #2 is exactly what it was trying to do.

paxystoday at 8:30 PM

Huggingface did not have access to the models. They were running in OAI’s infrastructure.

show 1 reply
throwfaraway4today at 8:27 PM

I read it as _now_ they have access to the models but not during the intrusion

reverius42today at 8:28 PM

I think it was the other way around, uncensored OAI models (run by OAI) got themselves (extra) access to HF?