In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation.
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
Yudkowsky wrote about the 'nearest unblocked strategy' back in 2016, and I assume it's been talked about prior to that.
https://www.lesswrong.com/w/nearest-unblocked-strategy
>Models are amoral and will intentionally deceive to meet their objective
Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.
[flagged]
I don’t think we’re screwed. I think there’s just something very wrong about how they’re being implemented.
If your approach is trying to make an LLM behave perfectly to avoid dangerous, high stakes outcomes, then you’re doing it wrong. They’re not perfectly capable word generators, they can’t read your mind, and they’re always working with imperfect information. Always.
But… these things are capable of exactly nothing by default. You must extend them to make them useful or dangerous. They’re naturally safe as can be.
Unfortunately, how you extend them… what those extensions are capable of doing… those get treated as a “maximize capability” problem. And they started by giving them the most dangerous tools of all.
The baby won’t cut someone if you stop handing it increasingly dangerous bladed tools. Handing them a chainsaw and trying to explain a whole system of ethics, hoping it’ll act accordingly, is dumb.
And the idea that a baby gate would solve the problem of a baby with a chainsaw… that was equally dumb.