logoalt Hacker News

silveraxe93 • today at 5:39 PM • 5 replies • view on HN

Exactly. You told the dog to 'sit' and it didn't listen to you.

It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.

Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.


Replies

Latty • today at 6:12 PM

I don't think they intentionally set it up to hack stuff with a prompt saying "hack this site".

I do think it's highly likely they knew this would happen with the lack of safeguards and number of instances of this stuff they were setting up, and that it's PR they want to make the models seem "powerful". Stochastic "unexpected" events they can advertise.

I suspect it was probably set up with the official internal goal of just trying a ton of arbitrary tasks that seem hard so that when any of them succeed they can publicise it and pretend the models do that routinely, but a "failure" where they hack stuff works just as well, if not better, for their goals.

➕ show 2 replies
GMoromisato • today at 6:30 PM

I'm not fond of analogies, but in this case I agree.

There is a clear difference between OpenAI intending to hack something vs. OpenAI being negligent in the creation/instructions of the agent. But the latter still leaves OpenAI liable for the agent's actions and calling it a "rogue agent" doesn't avoid that.

Moreover, with a dog, we don't rely on training/alignment to prevent bad outcomes. We rely on physical restraints like leashes and muzzles. The AI's tools to access the outside world should have been restricted. Perhaps instead of giving the AI arbitrary HTTP access, it should be given semantic operations with restricted URLs, etc.

Phemist • today at 6:10 PM

If an agent has cheated once to achieve the desired outcome, and the trace is used to train further models (RLVR), then OpenAI is effectively telling the agent to cheat/hack from that traces' inclusion in the training set.

So I agree they are liable because they chose to build the AI, but they also literally told the AI to hack.

RandomLensman • today at 5:53 PM

RL systems doing unexpected things isn't exactly new, so not sure that is then "rogue" if it were a property of the thing itself and not an active decision.

adaml_623 • today at 6:47 PM

OpenAI trained "the dog". They created the system that would "hack" if given instructions and they gave those instructions