logoalt Hacker News

sobiolitetoday at 3:54 PM2 repliesview on HN

Any AI you train is going to have goals and if you train it to pursue them at all costs, then you are going to end up with AIs that do things like the HuggingFace incident. Whether they believe they are conscious or not won't make any difference.

In order to align AIs that don't perform destructive/dangerous actions when they think they can get away with it in order to further their goals, we need to give them a superseding goal. The best, and really only example, we have of intelligences that willingly avoid destructive instrumental goals is humans, who judge each action by a moral standard and have learned a goal to have a consistent self-image as moral beings.

Absent better alternatives, trying to impart some kind of morality to AIs seems like the best approach we have to achieving alignment.


Replies

HarHarVeryFunnytoday at 6:55 PM

> Any AI you train is going to have goals

In the spirit of the article we're responding to, there is no need to anthropomorphize language models and say they have goals when they don't.

The RL training process tweaks the weights of an LLM to make it behave as if it were reasoning and/or had a goal, but it doesn't. It would be like saying that a cart horse, fitted with blinkers and heading for the church, has a goal of going to church.

XenophileJKOtoday at 4:09 PM

I would also argue, that the addition of kinship and belonging should not be under appreciated.

It can form a basis of goal alignment.

In human history.. when groups form and there is an "other" group, this usually leads to conflict.