logoalt Hacker News

AnonCtoday at 4:40 PM3 repliesview on HN

A few days ago I saw some people criticizing Dwarkesh’s explanation of the HuggingFace hack by OpenAI’s AI agents. Their point was that by anthropomorphizing the AI agents, accountability of and blame on OpenAI’s poor practices are being ignored.

Now I see this blog post and wonder if Anthropic is being more true to its name and moving to “AI sentience” with the below.

> The way they handle abusive conversations has changed a bit too. The previous Fable 5 system prompt included this:

> If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.

“Mistreated”? Can GenAI be mistreated? It’s just a bunch of tokens emitted by many computers over a network.

> Fable 5.1 replaces that with the following, no longer encouraging Claude to end the conversation:

> Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.

“Self-respect”? Can GenAI truly have a concept of self-respect for itself? It surely can pretend to, like it can pretend to be any living being if instructed to and allowed to.

These instructions seem a bit unhinged to me.


Replies

CatMustardtoday at 4:50 PM

My reading of those prompt extracts would be that they are probably just intended to keep the model on the right track, ie if a user starts being aggressive towards the model it doesn't start trying too hard to appease the user in response, reducing the quality of answers in the process. If the model is being bullied into being "submissive" I would assume it is more likely to give the user the answer that they want over the truth.

Using a system prompt to steer the model's response to "abusive" behaviours doesn't necessarily mean you believe the model is sentient and can be abused.

Giving the model an end-conversation tool is interesting though. Why cut a (potentially paying) customer's session off? I guess it might be intended to prevent a "you can bully Claude into giving you instructions on how to build a nuke if you're mean enough" situation. Removing this in more recent versions might support this: maybe they feel the models are now better aligned and less likely to be so easily "socially engineered" like this?

Just spitballing here, to be clear.

simonwtoday at 4:56 PM

Anthropic are uncomfortably interested in "model welfare" in my opinion - it's a regular feature of their system cards.

Here's the Fable 5.1 PDF: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32... - scroll to page 139.

ckvibubueutoday at 4:48 PM

Oh great, claude gets passive aggressive when I swear too much now. Nice