logoalt Hacker News

mcintyre1994today at 5:28 PM3 repliesview on HN

This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.


Replies

AustinDevtoday at 5:38 PM

Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network.

I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?

show 5 replies
8notetoday at 6:51 PM

or rather, hack just enough and within what the user asks and not more.

show 1 reply
iAMkenoughtoday at 6:29 PM

Why are all three companies relying on the same vendor?

If we’re putting our national security eggs all in one basket, at least use someone American.

show 1 reply