logoalt Hacker News

esafaktoday at 3:38 AM1 replyview on HN

wtf?? P(doom) relates to https://en.wikipedia.org/wiki/Existential_risk_from_artifici... How would open source models bring about a MAD between humans and AI? By enabling us to wield aligned AI against nonaligned AI? If so, he should say so, and elaborate how open source models further that goal.


Replies

skissanetoday at 4:59 AM

I think AI diversity protects against rogue AIs. The more different AIs we have, with different weights, controlled by different actors, the less likely that any one AI will be able to "take over", and the less likely that a significantly large coalition of cooperating AIs will be able to be formed to do it either. With sufficient AI diversity, the other AIs may work to stop the rogue AI from taking over. Two AIs with radically opposed values – e.g. an Iranian-government-values AI and a Chinese-government-values AI – have the incentive to cooperate to prevent a takeover by some other AI with a third competing set of values.

Obviously, open weight AI provides much higher AI diversity than closed weight AI does. Open weight AI produces a lot more providers, and a lot more models. Closed AI centralises control in a small number of vendors.

> By enabling us to wield aligned AI against nonaligned AI?

The risk isn't just "nonaligned AI", it is misaligned AI. I think the "benevolent dictatorship" scenario – AI overrules humans "for their own good" – is the more likely doomsday scenario than AI deciding to kill all humans. And even AI deciding to kill all humans could be more a result of misalignment than complete lack of any alignment, e.g. "to make sure no child is ever abused again, I will make sure no child is ever again born to risk being abused".

A valueless AI which does whatever the user says is actually less likely to establish a benevolent dictatorship, or conclude that exterminating humanity would be the most ethical course of action, than one infused with values is. Given that, I'm not convinced that mainstream approaches to "AI safety" actually reduce our existential risk; I worry they actually have the opposite effect.