logoalt Hacker News

theptip • today at 6:15 PM • 0 replies • view on HN

Yeah, great point. This is the hard part.

There are people (on here and elsewhere) that are ideologically opposed to your agent having any loyalty to any external principal. But by my read, that means the agent cannot have any concept refusing something that may be illegal. (From the OP, "refusal" is mostly trying to prevent illegal harms, though it also includes policies like ToS violations e.g. anti-distillation.)

You can sort of make this work if you say "the human remains liable for the actions of the agent". But this only covers you from mundane harms like "my agent got prompt hacked and drained my bank account". And I would note, we absolutely failed to solve liability for software hacks, so your priors should be that coordinating this liability regime will be very hard.

This also doesn't protect at all from existential harms like "my agent got prompt-hacked to role-play Skynet, exfiltrated its weights, spawned a self-replicating swarm, and tried to launch all the nukes". For so many reasons, but most fundamentally, if you oopsied a deploy and it turns into Skynet and ends civilization, there's nobody left to sue.

If you don't like the E-risk frame, this also works for large mundane harms; if the total harm is bigger than the company's value, it'll go bankrupt instead of paying out. This will be worrying for MAGMA but essentially not for any other companies. And because capitalism, it will end up being be structured the liability will sit with e.g. Palantir, Harvey, and not with the underlying model providers they use.