Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.
The AI is getting bad morals from listening to that dreadful rock and roll
Sounds just like the fantastical nonsense that comes out of Lesswrong.
They could filter what they train on if they wanted to - they just don't want to.
I mean, you're not wrong, but by that logic we were done for even before we had digital computers.
Related reading:
The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.
https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...