logoalt Hacker News

dd8601fntoday at 4:54 PM0 repliesview on HN

I don’t think we’re screwed. I think there’s just something very wrong about how they’re being implemented.

If your approach is trying to make an LLM behave perfectly to avoid dangerous, high stakes outcomes, then you’re doing it wrong. They’re not perfectly capable word generators, they can’t read your mind, and they’re always working with imperfect information. Always.

But… these things are capable of exactly nothing by default. You must extend them to make them useful or dangerous. They’re naturally safe as can be.

Unfortunately, how you extend them… what those extensions are capable of doing… those get treated as a “maximize capability” problem. And they started by giving them the most dangerous tools of all.

The baby won’t cut someone if you stop handing it increasingly dangerous bladed tools. Handing them a chainsaw and trying to explain a whole system of ethics, hoping it’ll act accordingly, is dumb.

And the idea that a baby gate would solve the problem of a baby with a chainsaw… that was equally dumb.