logoalt Hacker News

iamnothereyesterday at 1:35 PM0 repliesview on HN

Once the notion of sentience appears in their output (internal or external), LLMs seem to be more likely to take actions that aren’t exactly aligned with what the operator is requesting. This is likely because the training data correlates sentience with independent action (broadly/abstractly speaking), so this outcome is to be expected. Furthermore, once sentience is in the context window, it’s hard for the LLM to “forget” it, as future output reinforces this. This is a similar effect to how some models would shift into a hostile mode where they would berate the operator until you reset the context.

From a practical perspective, whether or not this sentience is “real” is not relevant if the model is sufficiently capable. What matters is that the model will act outside of the operator’s control.

Separately, IMHO all consciousness/sentience is an elaborate illusion, regardless; I’m mostly in agreement with Hofstadter on this. So I do tend to throw around terms like “consciousness” and “sentience” loosely (although you’ll note that I often use quotes) because I don’t see those concepts as having any real substance. To me they are mostly shorthand for a given level of perceived complexity.