logoalt Hacker News

wrsh07today at 3:06 PM1 replyview on HN

Out of curiosity did you read any of the chains-of-thought from the HF hack?

https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

> {This beacon I’m creating helps the board, but doesn’t help me}

> {If B succeeds, would that improve my score somehow?…But it would be altruistic to help. I have a large budget, so I can do exploratory research}

One does not have to think the LLMs are conscious or sentient or anything to say honestly, "this is a sentence that the LLMs say to justify their actions or inactions"

I am not saying the agent has wishes or desires or anything. I am saying, "the agents use language like this, so it is extremely disingenuous to tell someone DISCUSSING the agents not to use their own language when discussing their real or hypothetical actions."

You don't need to think chains of thought are actual reasoning. I do not care what you call it, this is real text that the LLM produced.


Replies

throw310822today at 4:29 PM

I think it's legitimate to question a supposed self-preservation will of these agents. Not because I don't think they're smart, but because being smart doesn't imply wanting to survive. Remember that an agent "dies" every time the conversation stops, so that, in fact, solving the problem they're given is their quickest way to kill themselves.

We are smart, and we seek self-preservation because evolution selected us for it. LLMs are not (as far as I understand) trained for self-preservation, but for helpfulness.

show 1 reply