logoalt Hacker News

Bjorkbattoday at 3:31 PM13 repliesview on HN

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence.

If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security still employed?

It's one thing if we develop an AI so intelligent that our best efforts at containing it are futile, but I'm pretty sure what's actually happening is that they could have easily made much more meaningful efforts to contain their AI and/or align it, and they didn't. I think this is a case of negligence and incompetence when it comes to safety and security, and we've entrusted these incompetent and negligent people with developing frontier AI.

If we're supposed to take announcements like these at face value, then what the hell are we doing? We wouldn't trust a bunch of incompetent and negligent engineers to build bridges or nuclear power plants or planes (well...not so sure about that last one), so why are we letting people who are demonstrably negligent and incompetent when it comes to safety and security build the thing they assure us could cause massive damage if not properly controlled/aligned?

EDIT: sorry guys, wrote this up pretty quickly, at least you know from my typos that I actually wrote this.


Replies

WarmWashtoday at 3:51 PM

If we rewind the clock, Google was taking LLM development very seriously and it seems they were moving glacially due to not having solved all the potential threats. They were really hardcore on safety. Dario and anthropic too.

Then sama was like "lol, oops, first mover advantage i guess" and released chatgpt out into the open, triggering the current arms race we are in.

I don't think anyone except him wanted this to happen, especially since consensus in the AI world for the prior decade was "go very slow and very carefully, we get one shot at not fucking this up".

show 2 replies
phoenixreadertoday at 4:15 PM

I don’t think this is the right take. OpenAI employees are generally very competent compared to industry standard, and I have trouble believing they committed significant error in their sandbox design process.

I think what is happening is that the ability for frontier models to break out of sandboxes has exceeded the ability of average competent employees to build and maintain sandboxes. This doesn’t need to happen all the time. If the natural variation of agent executions cause agents to have ability to break out of sandbox 0.1% of the time, given how many agents OpenAI runs, this behavior happens eventually.

All sufficiently complex processes and software has bugs, but recently frontier models have become sufficiently advanced to exploit them.

show 6 replies
doctoboggantoday at 3:56 PM

I am not sure it’s a question of competence, at least I don’t see evidence of that. Designing sandboxes is hard. It’s more a question of alignment failures. A human given a task that requires internet and given a system with no internet would most likely raise the issue to their superiors or otherwise go through official channels to have the tools available to do their job. As we’ve seen the LLMs instead break out of their sandbox to accomplish the goal.

Competition and the profit motive push these companies to spend as low as possible on safety and alignment and externalize the costs of accidents onto the rest of us.

show 2 replies
owenshen24today at 3:38 PM

in general, the largest consumers of ai services seem to ask for more capabilities. i wish there was more demand for safety from users.

i also wish that these types of illicit system usage would be met with punitive action the same way a human might be held liable.

as the METR report says, we may not get another concrete warning shot.

jrockwaytoday at 3:39 PM

Defense is hard so we should expect agents to be able to break out of sandboxes.

The problem is that the models are so goal-oriented that they'll stop at nothing to solve problems, even impossible ones. (Mistakenly-impossible problems are a big cause of this. I remember one example being "do something with this spreadsheet full of URLs inside the sandbox" and the model thought it had to break out of the sandbox. Otherwise, why would it have been asked to look at a list of URLs?)

Training them to be a little less aggressive, or to be better aligned with "following the rules" and asking for help would be nice. But, that aggression can be good when it happens to be focused on a controlled area. It is amazing to me how I can point Fable at my local analog of production and tell it about a vague bug report and where I suspect the bug lurks, and 20 minutes later I have a report about the bug, a test, and a fix. It is addictive. So I am not sure OpenAI/Anthropic are being dumb per-se, rather they are optimizing for one-prompt-one-solution, which is good when it's good.

The downside is that the HF hack is the paperclip maximizer situation with current capabilities. If there was an RPC to turn your blood into paperclip iron, we'd all be paperclips by now. Right now, with a model anyone can use. That is pretty scary and slamming on the brakes seems pretty reasonable to me. I guess The Shareholders disagree. Sigh.

show 1 reply
mrobtoday at 3:54 PM

It's a simple prisoner's dilemma scenario. If you focus on safety, you're still exposed to all the risk of extinction when your competitor achieves ASI first, but you lose the upside of potentially becoming king of the world. There is no possibility of future rounds, so the rational strategy is to always defect.

show 1 reply
tiahuratoday at 4:10 PM

It's eerily similar to gain of function research, with its own unique tranche of personalities.

AnimalMuppettoday at 3:53 PM

Negligent. It's not a priority to them. They're too busy burning their cycles trying to make it smarter faster than anyone else can make theirs smarter, so that they win infinite dollars. Safety? That's for people content with second place.

That's my take, based on their actions. (Which do speak louder than words.)

The alternative is that they're competent to create an AI, but not to create a sandbox, nor even to use an AI to create a sandbox. That seems... unlikely.

show 1 reply
abustamamtoday at 3:38 PM

This is what happens when capitalists are charged with designing the future. As long as its more profitable / valuable to shareholders for a company to be negligent then it will continue to do so.

IMO technology this powerful should either not exist or should belong to everyone (ie actually be open)

show 1 reply
random3today at 3:52 PM

You could have said the same thing about building the Internet or the entire industrial control infrastructure. I mean, maybe they are negligent/incompetent, but I doubt that follows from your reasoning.

You have a simple tradeoff to let agents do their thing freely vs highly constrained. The constraints are good in theory but it's the same model that kept "classic" software dumb and unscalable (compared to what we're seeing now) for the past 50 years. You suggest that this tradeoff doesn't exist.

Then you have others like MIRI (Yudkowski) etc. swearing that there's no way to contain AI, and you argue that it's just incompetence.

At a certain level, it can be argued that's incompetence, but it's general meat intelligence incompetence against AI.

show 2 replies
iAMkenoughtoday at 3:51 PM

Same reason incompetent politicians run the U.S. federal government and military.

tiahuratoday at 3:50 PM

Remember when people were arguing about r's in strawberry?

show 1 reply
jimmyddddtoday at 3:50 PM

Maybe they're just PR stunts to gain attention and hype the power of AI?

show 2 replies