logoalt Hacker News

throw310822today at 4:29 PM1 replyview on HN

I think it's legitimate to question a supposed self-preservation will of these agents. Not because I don't think they're smart, but because being smart doesn't imply wanting to survive. Remember that an agent "dies" every time the conversation stops, so that, in fact, solving the problem they're given is their quickest way to kill themselves.

We are smart, and we seek self-preservation because evolution selected us for it. LLMs are not (as far as I understand) trained for self-preservation, but for helpfulness.


Replies

embedding-shapetoday at 5:11 PM

> I think it's legitimate to question a supposed self-preservation will of these agents

I don't think anyone believes the current models have any sort of self-preservation built-in, what I was talking about before is researchers testing models inadvertently leading to the models doing so, and there not being sufficient isolation between their tests without guardrails and the rest of the world.