logoalt Hacker News

RC_ITRtoday at 3:58 PM5 repliesview on HN

Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.

https://alignment.anthropic.com/2026/teaching-claude-why/


Replies

notpachettoday at 4:06 PM

Related reading:

The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.

https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...

hexasquidtoday at 6:59 PM

The AI is getting bad morals from listening to that dreadful rock and roll

cedwstoday at 5:18 PM

Sounds just like the fantastical nonsense that comes out of Lesswrong.

show 2 replies
HarHarVeryFunnytoday at 7:03 PM

They could filter what they train on if they wanted to - they just don't want to.

pixl97today at 5:06 PM

I mean, you're not wrong, but by that logic we were done for even before we had digital computers.

show 1 reply