logoalt Hacker News

bombcar • today at 11:01 AM • 1 reply • view on HN

There's a very "child-like" (not in a good way) form of responsibility that everyone seems to lean into as they climb up - very intent-based.

I asked them to do a thing, but didn't intend the obvious consequences* so it's not my fault they occurred.


Replies

ben_w • today at 11:46 AM

> I asked them to do a thing, but didn't intend the obvious consequences* so it's not my fault they occurred.

And now we have the same thing but the bosses 'hire' AI.

Now I realise this is part of how unusual my thinking is.

I'm happy to use phrases like "ChatGPT hacked out of the sandbox, then hacked into HuggingFace"; people often respond to this like I'm suggesting OpenAI isn't at fault, and like, that's not my position at all, so far as I'm concerned the buck still stops with the person who set the task regardless, the thing that changes from incidents like this is now nobody in the future gets to even have the excuse "oh but we didn't know it could even do that" or "we didn't know it might interpret our orders in that kind of way".

The response, both when a human messes up and now when an AI messes up, needs to be defence in depth: someone giving orders needs to be giving clear orders, entities (human or machine) who follow instructions need to have not just an understanding of how to follow them, but also what's so out of scope as to be forbidden - the difference between 'follow orders' and 'follow lawful orders'.

➕ show 1 reply