logoalt Hacker News

reasonablekloutyesterday at 11:14 PM2 repliesview on HN

We also now have observed multiple major incidents where AI agents behaved in completely unanticipated ways (forming a collective, hacking their own eval infrastructure) and attacked public infrastructure without being told to do so, without any of the human developers noticing.

Even if "AI will cause human extinction" is still unclear, we have plenty of proof that catastrophic damage is possible, the industry is developing the technology in a reckless manner and that all the hypothetical safeguards ("we can just pull the plug", etc.) are simply not present today.


Replies

atherton94027today at 12:07 AM

If you've run an ssh server connected to the internet you've gotten used to the hundred of bots probing it every day. How is this AI threat different from humans writing scripts to pop linux servers?

lostdogyesterday at 11:41 PM

Those behaviors have been anticipated for years if not decades. And each incident is minor and leads to clearer rules and safety behaviors for AI agents.

show 1 reply