There is one thing that is still unclear to me after reading a lot about the hack:
* were the ExploitGym solutions actually available somewhere inside huggingface's private datasets ?
* was the model really trying to extract the solutions ? or had some sub-agent drifted enough from the original context that it was not even trying to solve the initial challenge ? that would look much worse for OpenAI, PR-wise.
According to https://cloudsecurityalliance.org/artifacts/hugging-face-cis... the models ended up finding CyberGym solutions, which was the wrong benchmark.