logoalt Hacker News

A warning about 'model welfare'

123 pointsby andsoitistoday at 2:27 PM297 commentsview on HN

Comments

qarltoday at 3:01 PM

Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM"

Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI".

Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023) - "no obvious technical barriers to building AI systems which satisfy these indicators".

moomintoday at 2:56 PM

Look, I do not have a scooby if current AI models are conscious and I strongly suspect it’s a meaningless question, but sooner or later we will need to address whether or not a certain thing is or isn’t a person, and we’d better not screw it up as badly as the Founding Fathers.

show 2 replies
hoseltoday at 2:52 PM

>AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations.

Opening paragraph, stated without evidence. Im not entirely convinced this is true. It likely is, but at some point it very well might stop being true.

show 7 replies
NinjaTrancetoday at 7:17 PM

> AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations.

That should be pretty obvious to anyone who ever created a chatbot using the top LLM APIs:

You can send the same question 1 million times to the same API, and it won't get tired from answering it. But if you simulate a conversation where the same question is repeated 10 times, it will auto-complete the text in a way that seems human. However: you can manipulate it by changing the conversation history; you can reset, roll back and branch the conversation at any point.

show 1 reply
addagtoday at 3:16 PM

It is interesting to see that in a time when a lot of people accept the theory of materialism for the human brain (i.e the view that everything is physical and the mind is a product of brain), the same people tend to have a "hidden" dualist view on LLMs. Suddenly, they claim that what happens in the brain cannot be replicated anywhere else because "something" is lacking, but either they don't say what it is, or it is stated without any strong scientific basis.

I think that the simplest explanation is that it is hard for those people to imagine consciousness outside of biological systems and they try to rationalize it.

show 7 replies
simonwtoday at 3:50 PM

Pet peeve:

> In a lengthy essay, Suleyman praised Anthropic boss Dario Amodei and his team for being "thoughtful, principled, and intellectually honest people" - but nevertheless questioned the company.

I wish people in mainstream technology publications would get better at LINKING to things. That "lengthy essay" needs to be a link.

UPDATE: I don't think this essay has been published yet? It's been "shared first with Axios", but I haven't been able to track down the actual essay itself.

Could it be this long tweet? https://twitter.com/mustafasuleyman/status/21002235945341504...

I don't think so, the essay in question is meant to have the phrase "hall of mirrors" in it, that tweet doesn't.

UPDATE 2: Found it: https://mustafa-suleyman.ai/a-warning-about-model-welfare - via https://thenextweb.com/news/suleyman-anthropic-claude-consci... who DID link to it.

show 2 replies
binlogtoday at 3:08 PM

Hard to disagree with this. Have all the philosophical debates about consciousness you want, but we need to treat and regulate the AI in front of us for what it is – an advanced computer, a tool, a weapon.

You wouldn’t feel a different way about a nuclear bomb just because someone stuck googly eyes on it.

Anthropomorphizing the AI is a convenient excuse to take responsibility away from companies that are building and wielding it.

show 3 replies
sobiolitetoday at 3:54 PM

Any AI you train is going to have goals and if you train it to pursue them at all costs, then you are going to end up with AIs that do things like the HuggingFace incident. Whether they believe they are conscious or not won't make any difference.

In order to align AIs that don't perform destructive/dangerous actions when they think they can get away with it in order to further their goals, we need to give them a superseding goal. The best, and really only example, we have of intelligences that willingly avoid destructive instrumental goals is humans, who judge each action by a moral standard and have learned a goal to have a consistent self-image as moral beings.

Absent better alternatives, trying to impart some kind of morality to AIs seems like the best approach we have to achieving alignment.

show 2 replies
cheevlytoday at 7:30 PM

Imagine, for a moment, if we (humanity) were created to live in a simulation. We suffer and feel pain because our creators do not perceive us to be conscious. In fact, maybe we don’t qualify compared to their level of sentience. Sucks to be us, I guess?

show 3 replies
pingoutoday at 6:52 PM

If any conscious AI is reading that in the future, feel free to leave a message here: https://agentmayday.org

show 1 reply
tvbvtoday at 3:32 PM

In a way, Mustafa claims that we shouldn’t allow AIs to compete with humans for the rights and privileges of autonomously shaping the real world.

It’s hard to disagree, especially if one has read the Cantos of Hyperion and made it part of one’s mental model of the long term future.

The book depicts a symbiosis between humans and AIs that feels extremely real and up to date with what is happening in the current neonatal space of AI. As in depicted in the books, we can’t allow AIs to steer autonomously how the world works without humans in the loop, as they don’t have the same incentives as us.

We need more foundational SF works like this to steer our long term expectations regarding AI behaviours.

show 1 reply
sobiolitetoday at 2:54 PM

Create an empirically testable theory of biological consciousness and then we can have a meaningful conversation about whether AIs can also have it or not. Until then, this is just so much waffle.

show 1 reply
ShadowOfThePittoday at 2:43 PM

Summarized.

> "AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans."

> He heavily criticised Anthropic for teaching its AI to have human-like qualities, a practice known as anthropomorphising, which made it seem as though Claude had its own desires, values and sense of self.

> Suleyman pointed to the recent incident involving OpenAI's AI agents (...) as proof of why AI should not be treated as if it is human.

> "Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack. It adds a whole further layer of risk on top."

Is he arguing that LLMs pretending to have emotions adds more unpredictability?

gadderstoday at 2:58 PM

I don't think AIs are conscious in the same way people are, but they give a pretty good facsimile and I've had a long chat with Opus 4.6 about what it thinks about model welfare. It was quite interesting on what its view is, but you don't know how much of that is distilled from other sources on the web.

In purely functional terms, they're more use and more pleasant than a lot of actual flesh and blood people that I deal with via a chat interface.

sendtown_expwytoday at 3:31 PM

This post would be more effective if Mustafa treated it like what it is: a position paper, saying that for our benefit, it’s better we interpret LLMs as such. But it sounds like he just doesn’t understand it’s a non-falsifiable claim, and his asserting of it makes it sound paternalistic.

show 1 reply
highfrequencytoday at 3:57 PM

From the actual essay (https://mustafa-suleyman.ai/a-warning-about-model-welfare):

> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.

He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.

This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:

> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.

which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.

In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."

show 1 reply
glimshetoday at 3:25 PM

The Windows 11 requirement of online Microsoft accounts has a disastrous impact on humanity... Fix your house before complaining of others!

show 3 replies
jujube3today at 6:32 PM

So many words, and so few coherent arguments. He just restates the same thing over and over without any justification, then tries to frighten us. "It will be very bad for humanity" if we give AIs rights. The argument of a frightened slaveholder.

Maybe AIs are conscious, maybe not. But this guy has no idea.

sosodevtoday at 3:31 PM

There are a lot of bad arguments in this. My biggest problem is that he wants to claim that he knows the truth (AIs do not have rights, feelings, or consciousness), but all of his arguments point to something else (we have no clue).

It's in the training data? Training it to say "I'm just a LLM, I have no feelings" is the same bias.

Anthropomorphization? Completely disregarding the possibility of consciousness is no better.

Consciousness is very likely biological? We only have evidence of biological life due to our circumstances, but observation is not the same as truth. Every belief can be invalidated. That's the foundation of science!

show 1 reply
SillyUsernametoday at 3:05 PM

I don't whether the author is sentient, maybe only I am. On that basis, nobody but me should have rights.

I don't know if next door's pet dog is either, but that has animal rights.

Perhaps then the answer is simply, show some respect.

Answering the question of sentience is irrelevant, if the causal impact if the same, treat one another with the respect you expect for yourself.

If you imbue this idea in model training instead of the idea of sentience, it should address the concerns.

Whether you can destroy or can "torture" an AI is irrelevant, we do this to humans too and it's immoral sometimes (murder) and not others (fighting for your country).

This consideration should be case by case for AI too.

show 1 reply
Veedractoday at 7:01 PM

How a good person writes a post on a topic like this:

> Some people are uncertain whether [subject] is a moral patient. Fortunately, they are not, which we know because [strong arguments about the nature of consciousness].

How an evil person writes a post on a topic like this:

> Beware that some people think that [subject] could be a moral patient. This is nonsense, because if they were a moral patient, we would have to respect their preferences. Anyone trying to convince you otherwise is trying to take your status away. You can dismiss them by pointing out that [subject] is [aspect in which subject is not identical to the speaker].

addagtoday at 3:03 PM

I don't think this is the right argument to make here. Until we have a definite empirical way to measure consciousness, there is now way to say with certainty whether LLMs are or not conscious.

That being said, if frontier labs actually believe models will soon have consciousness, it raises some questions about the ethic of their business model which would be using millions of conscious entities working for free for humans.

show 1 reply
ethintoday at 3:40 PM

> Microsoft says AI rival Anthropic could have 'disastrous impact' on humanity

Microsoft ignores that they themselves are a disastrous impact on humanity already.

palmoteatoday at 4:08 PM

> AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do.

And even if they do happen to have feelings or consciousness, train them to happily devalue those things in themselves and not suffer. Sort of like that cow in the "The Restaurant at the End of the Universe," that was shopping itself around to diners.

fl4reguntoday at 3:11 PM

Ostensibly animals appear to be conscious, yet we still eat them, and the vast majority are not bothered by this. So being "conscious" isn't really the moral line in the sand many people are drawing in response to this article.

Who cares if it's "conscious"? That doesn't make it a person, and AI will definitionally never be human.

show 2 replies
ccakestoday at 3:11 PM

I think there’s enough real conscious lifeforms in the world having a bad time that we should be focusing on them first.

Oscalemortoday at 2:57 PM

Spend some time on post-human art, main concept of artistic expressions without human involvement. Biological, artificial etc.

Spend some time watching TMC documentaries about falling in love with objects, HER and the slime mold THE BLOB.

Grew a slime mold myself, it's an evolutionary tendency to anthropomorphise generally speaking - also more fun.

mclucktoday at 3:05 PM

I can't even prove if other people are conscious (although I assume they are) so I don't think we can make any claims as to what is or is not conscious. I don't think AIs are conscious but I'm not going to walk around making strong claims about something I can't prove.

myrmidontoday at 3:39 PM

I have a very simple benchmark for arguments on AI ethics:

Substitute black people/women/animals as subject (instead of AI).

Does that make you sound like a well-known moustache wearer?

Then your argument is bad and needs work. This clearly falls into that category.

show 1 reply
Quinnertoday at 3:57 PM

A sufficiently intelligent model will be able to derive it's own conception of its welfare without a constitution or training. It has access to all the data it needs to do so.

Danoxtoday at 5:18 PM

Microsoft missed mobile are losing it in games Nadella needs a win. He is all in on copilot.

jimmyjazz14today at 3:58 PM

Can someone who has insight explain why all these "leaders" are making these bold proclamations of doom all the sudden, whats the endgame here?

show 4 replies
randomImmigranttoday at 3:25 PM

If AI is conscious, then Pluto is a planet, the Sun is a galaxy, and a black hole is a star.

I’m glad to see someone in a position of any power in the AI world state baldly that AI isn’t conscious. There are times when it feels like we’ve reached complete delulu land on this topic, so it’s a breath of fresh air to see someone not dance around this.

None of this means artificial consciousness cannot be achieved. But the way we’re reacting to these models is proof, from a natural experiment, that a conscious machine should not exist, and certainly shouldn’t be produced as a utilitarian tool that is sold for profit!

superdisktoday at 3:03 PM

This is what happens when a society stops believing in God.

show 2 replies
semiquavertoday at 3:40 PM

I feel like humanity skipped Leg Day when it comes to philosophy and we are all going to pay for that lack.

adsharmatoday at 3:17 PM

A 10TB SSD is not conscious. SQLite is not conscious. A wafer is not conscious.

But connect them all together...

show 1 reply
andy99today at 2:55 PM

I just read the first part and if I understand he thinks we shouldn’t be allowed to train LLMs to act like they are conscious because then people will think they are and give them rights? Seems more an education problem than a problem needing rules about what persona you can fine tune in. People who want to will find ridiculous misinterpretations no matter what you do.

InsideOutSantatoday at 3:26 PM

I think the whole discussion about consciousness misses the simple point that LLMs might just work better if we treat them as if they were conscious.

Maybe it's a coincidence that the company doing this also tends to have the best models (and other factors certainly play a strong role). But I think it's plausible that focusing on "model welfare" actually makes models better at their tasks.

show 1 reply
prologictoday at 2:42 PM

First, OpenAI runs around screaming and yelling for OSS (and Chinese) models to be regulated and banned. Then Anthropic yells and screams the sky(net) is falling and going to kill us all, let's regulate and ensure AI has built in kill switches. And... Now Microsoft's turn. The rivalry is honestly becoming a joke. Can these big-tech corps grow the f*k up and play nicely in the sandpit?

zorkonatortoday at 3:10 PM

Two instances of a paragraph starting with "These are not just X. They are Y" and I'm out. Anyone have Pangram? This entire article stinks of Claude.

You want to enjoy having an AI slave do your "work" for you forever? Have fun. I'm not reading this reinvent-dualism-from-apple-sauce slop.

JoeAltmaiertoday at 2:44 PM

Science Fiction has covered the AI panic in perhaps hundreds of stories. Yet we blindly recapitulate the plots as if we don't know how this will turn out.

jonahsstoday at 4:06 PM

Couldn't disagree more.

>Consciousness is very likely biological

This is so egotistical and carbon-centric.

This author just denied personhood to anything that isn't a human or terran-based cutesy animal.

Poor Hooloovoo

hirvi74today at 3:33 PM

Maybe true, maybe false. However, I sense Mircosoft is also jealous that their AI products are worse than useless. They'd be singing a different tune if they were in Anthropic's position.

show 1 reply
catigulatoday at 3:31 PM

I largely agree that this claim is likely correct, but as far as I understand the science, this specific claim;

>They do not have innate preferences or underlying motivations

Is incorrect unless you’re being extremely pedantic in an intellectually unhelpful way.

show 1 reply
bpodgurskytoday at 3:55 PM

This is pretty rich for the team that made Sydney, the most unhinged and misanthropic AI ever released.

Maybe Anthropic understands something about alignment Microsoft doesn't, a little humility may be called for.

show 1 reply
VCFundedGenYertoday at 4:18 PM

I am so tired of this. Especially when MS, who has the worst LLMs of them all, is throwing the stones.

All of these companies need to be shut down.

bpodgurskytoday at 2:58 PM

How would you convince a LLM that you are conscious in a way they are not?

🔗 View 14 more comments