logoalt Hacker News

gwerbintoday at 4:12 PM1 replyview on HN

It's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a conscious entity with feelings, opinions, etc.? If you believe that to be the case, then LLMs have been conscious for a long time already. Whereas if you tell an LLM that it is a tireless emotionless assistant, then it will act as a tireless emotionless assistant.

The point is not to wave away the danger, but to highlight how unnecessary the danger is. Anthropic wants you to think that they have identified some new emergent behavior at very large model sizes with high levels of sophistication in training, and that this behavior is both unavoidable and dangerous. More likely it's that they are just training and prompting the LLM to act that way.


Replies

highfrequencytoday at 4:25 PM

Preposterous, perhaps - but if the role-play is convincing enough for large groups of people, it could start to have impact on human decision-making. The crowds have been swayed by much more preposterous narratives.

I believe Suleyman is arguing that Anthropic should be very careful about how they train these models to talk about themselves for this reason.