This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.
or rather, hack just enough and within what the user asks and not more.
Why are all three companies relying on the same vendor?
If we’re putting our national security eggs all in one basket, at least use someone American.
Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network.
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?