Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network.
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
This was not the case for the hugging face hacks, as in those the agents hacked hugging face specifically on purpose, and they were trying to mask commands indicating they were had breached the "sandbox".
I mean, I get your thought process and don't disagree. That said...
Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!"
The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.
So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.
> Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events.
As I've said before on this website, fool me once on this.
If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
[dead]
I'm not familiar with the hacks this article is actually referring to, but I don't see how the HuggingFace attack could have worked based on that premise. They knew they had internet access, they knew they had working credentials for HF, they knew they were uploading malicious files, they knew they were trying to open PRs that HF would review. You obviously could build a simulator with fake HF infrastructure, but I'm not aware of any evidence that's what they thought they were attacking in that case.