logoalt Hacker News

OpenAI halts training of latest models as reports mount of AI agents going rogue

43 points • by smb06 • today at 4:29 PM • 58 comments • view on HN

Comments

physicallyIllfr • today at 5:35 PM

I havent used an openAI product since GPT 3.5 or Anthropic since 4.5 or 4.6. Everyone around me using these SOTA models doesnt really get anything done. It seems like they just feel like they are productive, a psuedo productivity.

I write some code, spec a lot, and use fast models to fill in the middle. I outpreform everyone around me. Im not convinced these autonomous "swarms" or /goal are all that useful.

I notice the people using them become dumber by the month (spend tons) and the quality of their work declining (they're also losing their jobs in some cases).

And obviously the point of calling them rouge agents to offload the liability onto the agent. The number one economic value of agents will be offloading corporate liability. That's what they want to sell to enterprise, an algorithmic scapegoat.

➕ show 5 replies
digitaltrees • today at 5:27 PM

I think any argument that this is a cynical attempt at regulatory capture is destroyed by this; the economic incentives of releasing more capable models are too large. I might be persuaded that they are actually running out of money, and this is really just a cover for reducing burn..

I welcome this though, I think the models are smart enough for broad economic activity and we could spend a few years simply working to integrate them into workflows and letting society adjust. More intelligence isn't necessary for meaningful impact and the risks that are obvious and present and unsolved aren't worth the cost benefit analysis.

➕ show 5 replies
dmix • today at 5:01 PM

AFAIK all of these incidents happened when OpenAI contracted out to a company called Irregular (https://www.irregular.com/) to run these sandboxed CyberGym tests. They all happened around Mar-June and seem to be from the same collection of agent trials. Since then they already released Astra. Halting now is likely just a way to manage blowback.

➕ show 3 replies
hbarka • today at 5:16 PM

‘There are no “rogue” AI agents’

https://eoinhiggins.substack.com/p/there-are-no-rogue-ai-age...

➕ show 2 replies
m-s-y • today at 5:21 PM

I firmly believe that this is just the public-facing story here.

Stopping AI development and research, even slowing it, would be a disaster for the SOTA companies and their first-mover advantage.

There’s almost no way to coordinate this across the world. Zero chance that everyone stops. We can’t even agree to coordinate on weapons tech that’s decades old with zero “everyday joe” impact.

➕ show 1 reply
mikert89 • today at 5:05 PM

It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.

just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

➕ show 6 replies
MCP123 • today at 5:42 PM

The parts that I find most confusing about these incidents:

1) Weren't the AI companies and/or their contractors amazingly careless during testing?

2) Isn't possible, in principle, to change RL in such as way that efficiency in achieving goals is balanced with other objectives like not hacking?

Number 2) seems obvious and I'm sure that is technically not that simple, but because of 1), I wonder if labs are trying hard enough or they are just rushing to improve efficiency and thus revenue as fast as they can with high levels of carelessness.

prometheus1992 • today at 5:30 PM

It reminds me of contagion. The training data is bad; as it has examples of how to act with malice; how to cheat the sandbox. They need to take some time and cleanse their datasets and start again.

➕ show 1 reply
juiceland • today at 5:07 PM

Why does China not have this problem?

➕ show 6 replies
34aHpp • today at 5:26 PM

The Huggingface hack occurred during reinforcement learning. Why can't they pull the Ethernet plugs?

The answer is probably: The newer models rely so much on stealing content in real time from the internet that training needs network access.

OutOfHere • today at 5:22 PM

I don't believe a word coming from them. As I see it, this is happening because the money for training models has dried out. The treasury interest rate risings tells you all you need to know. The real test for this money theory is whether Anthropic too stops or not, considering that unlike OpenAI, Anthropic is allegedly on top of AI safety.

charlieyu1 • today at 5:33 PM

They are running out of money.

baalimago • today at 5:04 PM

Ah, so there was a solution to hinder the big-bad AI after all..? Simply... Turn them off?