logoalt Hacker News

AI companies in race to demonstrate their model most threatening to humanity

338 points • by ljewalsh • today at 8:35 AM • 260 comments • view on HN

Comments

Spacecosmonaut • today at 1:48 PM

My read is that OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized due to various risks. E.g., an engineer deploys an agent over the weekend that decides, when stuck on a task, to go about hacking a competitor. They have a product liability issue.

It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.

They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.

The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

➕ show 11 replies
chasd00 • today at 11:42 AM

I’ve never seen CEOs work so hard to make the public aware of how dangerous and out of control their flagship product is. It makes me automatically assume they’re scheming about something else like regulatory capture to protect their market.

➕ show 29 replies
runako • today at 12:40 PM

Because the downsides that exist for other companies simply do not exist for this tier of rich companies. Examples:

- product liability

- negligence (civil or criminal)

- Computer Fraud & Abuse Act (requires intent, which after N "accidents" seems like a jury should at least evaluate whether intent is present as understood in a courtroom. Hard to blame "surprise" after the Nth "accidental" breakout.)

The bottom line is that if you or I trained a local model and it did any of this stuff, we would experience Consequences. ("Don't try this at home!") But an artifact of our unequal legal regime is that big rich companies generally do not and thus brazenly touting their immunity is part of their business strategy.

➕ show 2 replies
stephbook • today at 12:49 PM

We've had movies like Terminator, Matrix, or i, Robot. Black Mirror Metalhead. Boston dynamics robo dog with a knife or a gun. Humanoids doing back flips. Memes making fun of how we are going to fight these for access to water.

The intelligence is there and so is the physical incarnation.

Google's models fold proteins better than humans and Covid might have been engineered and inadvertently escaped from a lab.

But when CEOs warn of that, they're dismissed as scare mongers. Why? Are only powerless internet commentators allowed to be fearful?

➕ show 4 replies
xbmcuser • today at 12:56 PM

To me, the funniest part is that they pretend reaching superintelligence will somehow fix everything and make them win. But from what I can see, even if they do reach it, then what? What will that actually do? They don't have the manufacturing capacity to even use that intelligence. It will take decades to go from superintelligence to having the manufacturing capacity and the hardware, like robots, needed to make it truly useful. If I were China, as soon as the US reached superintelligence, I'd ban all exports of robots and the raw materials to build them.

➕ show 4 replies
ACCount39 • today at 11:38 AM

Because AI genuinely is an extremely powerful and extremely dangerous technology, and the "best practices" of dealing with that are still being written.

OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.

➕ show 6 replies
TuringTourist • today at 11:12 AM

Ah, I see we have now moved on to proving who has the tormentiest nexus. Never underestimate humankind's ability to outshine its own hyperbole.

➕ show 1 reply
someonebaggy • today at 12:15 PM

Just read the tone of this article:

> said Dr. Andrew Lenson, a Senior Lecturer in all kinds of science sounding stuff at Victoria University.

Don't we need more humorous yet serious writing like this in the world.

amelius • today at 12:24 PM

Why aren't the AI companies hacking each other?

Then we can truly see who is more dangerous?

:)

And I'm waiting for the story of Amodei sending a 6ft humanoid robot to OpenAI headquarters asking for Sam Altman.

heisenbit • today at 11:47 AM

Of course it must be dangerous - how else would you pitch it to the military?

christophilus • today at 11:49 AM

We are living in a timeline written by The Onion.

➕ show 2 replies
cmiles8 • today at 12:10 PM

Anything to take attention off the fact that these companies are burnings piles of cash and questionable future spend commitments with no sign of that stopping any time soon.

➕ show 1 reply
amelius • today at 1:25 PM

If I were the US government I would force the AI labs to solve the jobs that nobody likes to do, like cleaning the streets and collecting garbage. Because those are the jobs we will be doing if developments continue as they have been.

➕ show 1 reply
ks2048 • today at 12:10 PM

A recent Saturday Night Live clip featuring an impression of Amodei that hits on this theme,

https://www.youtube.com/watch?v=-Nvne3LzBls

sehw • today at 1:17 PM

Can they demonstrate that this crap actually makes money?

➕ show 2 replies
pizza234 • today at 12:50 PM

Stop posting garbage.

> Australian Prime Minister Anthony Albanese spoke with Altman to express “extreme concern” about the incident, a compliment Altman said he greatly appreciated.

"a compliment Altman said he greatly appreciated" is fabricated; there is no source for this.

➕ show 2 replies
smetannik • today at 1:29 PM

If this approach works with AI companies, why the same doesn't work for nuclear power plants? :)

solenoid0937 • today at 1:17 PM

I remember when HN scoffed at the idea of Mythos hacking anything.

a3w • today at 1:46 PM

Somehow, not the Onion.

MotoriX • today at 12:37 PM

At this point, AI companies are basically competing over who can destroy humanity first. What a time to be alive.

willtemperley • today at 12:33 PM

I had to double check if this was an Onion article or not.

Oh dear. I appear to still be stuck in the onionverse.

Kuyawa • today at 1:20 PM

Manipulation through fear

They need to control, regulate and tax AI so the first line of attack is the sheeple, so tell them the water is going to disappear and everybody will die a slow death of anguish, thirst and drought like never seen. Then evoke terminator apocalypses where drones go rogue and kill half the world. Then send their most likable emissaries with puppy eyes asking for compassion about our children's future

Then they ask for your vote (oh democracy)

In the end all they want is control and money, while keeping their competitors (china) in check, backed by the outrage of the weak majority

Politics as usual, move along

➕ show 1 reply
christkv • today at 12:32 PM

I wish they would call their bluff. Just declare them a national security concern and send in the feds. Lets see how fast they back peddle on this. Any if their models are hacking everything like they claim they should be charged for it like any normie would be if they did the same.

➕ show 1 reply
gitghxst • today at 11:58 AM

this is good for them because they can call for regulation and since they have infinite money, they can survive regulation and hurt labs that do not have infinite funding / open weights ones

sscaryterry • today at 12:23 PM

Honestly, I'm running out of ways to express what Sam and Dario are. I'm now down to 1 word: Wankers.

➕ show 1 reply
esafak • today at 1:33 PM

It amuses me to think that there could a noble aim: they're overdoing it partly to raise awareness of the real dangers. It's something we get as a side effect of their desire to achieve regulatory capture.

amelius • today at 11:43 AM

Meanwhile, users don't seem to care one iota.

ff10 • today at 1:13 PM

Imagine if the atom bomb was created by private companies about to make their IPO.

mminer237 • today at 12:48 PM

It's just marketing. They want people to believe they're selling super intelligence that is powerful enough to destroy the world of it fell into the wrong hands. It's much more effective than saying that you have a slightly more effective slop generator which obviously has no actual intelligence.

➕ show 1 reply
mpalmer • today at 12:55 PM

FWIW, the site's "quote of the week" looks entirely nonsensical and made up.

buellerbueller • today at 1:59 PM

If people had known about a civilian, peacetime Manhattan Project and were in a modern media environment where we could kibitz about it, would it have been slowed down?

Is AI/AIG too powerful to be undertaken by entities that are not beholden to our constitution? (I know that doesn't seem like much in the Trump era, but I hope for a return to government norms relatively soon.)

tantalor • today at 12:02 PM

Why is this written in the same tone as an article on The Onion? Is this pure satire?

➕ show 2 replies
petesergeant • today at 12:05 PM

"These models are harmless and it's all just marketing" is the HN equivalent of Covid-truthing. I was responding to someone on Bluesky recently who claimed the METR report said the whole coordination thing was hallucinated by agents reading the logs. Hadn't read it themselves of course, and were taking tiny fragments out of context to support it. These people seem to either believe that there's not really been any malicious agent access, or that all systems that can wreak havoc on us are perfectly air-gapped. I really can't wrap my head around it.

➕ show 4 replies
Gareth321 • today at 11:55 AM

This is the funniest timeline. Companies begging for heavy handed regulations because they promise their products will destroy all life on Earth. Governments absolutely refusing to govern. CEO of world's most valuable chip company interviewing in his Fonzie leather jacket telling people to BUY MORE OF MY PRODUCT BECAUSE THOSE SCIENTISTS DON'T KNOW SHIT ABOUT SHIT! World's richest capitalist promising we will have communism in 10 years.

shevy-java • today at 2:11 PM

I do not need a model to know that selfish mega-corporations cause a lot of harm.

Can we get our money back here? I mean, just take the increase in RAM prices alone. That should come with a huge penalty they should give us back due to collusion. And that's just one of many more things. Trump's family dynasty also owes us all money - we want that back as well.

everyone • today at 11:55 AM

pretend, not demonstrate

Arodex • today at 11:33 AM

Seriously:

Domyn's CEO says OpenAI, Anthropic are lying about safety

https://news.ycombinator.com/item?id=49875725

➕ show 1 reply
cfiggers • today at 12:33 PM

Consider this "strong" premise:

    So-branded "AI" models (that is, big bags of numbers and the linear algebra needed to exercise them) can be inherently or autonomously dangerous at a scale that threatens the long-term survival of the human race.
Consider also this "weaker" form of the same argument:

    So-branded "AI" models can be inherently or autonomously dangerous, if not at a scale that threatens the long-term survival of the human race, then at least at a level that significantly threatens, inconveniences, or harms humankind to a noteworthy extent.
Here's a few hard truths that no one will like to hear:

1. Those who accept and advocate for either or both of these premises do so based on ideology, and not based on the scientific method.

This is not a controversial statement. Advocates of the "AI danger" premise (whether the strong or weak forms above) freely and gladly, even insistently, note that the conclusive proof of what they're describing has never been observed, much less observed enough for any kind of experimentation loop to have run for any meaningful number of iterations. They are generally fine with accepting the premises without conclusive proof, because, they theorize, the instant that a conclusive proof event occurs and we observe it for the first time, humanity goes extinct shortly after. So if either premise is true, the consequences of them being true can only be averted by accepting the possibility / probability / certainty of them being true and acting to prevent their consequences without waiting for conclusive proof.

2. Statement 1, which asserts a brute fact and not my or anyone else's opinion, is not a commentary on whether either premise is in fact true or not. It is only a commentary on the epistemics of those who broadly accept them.

3. Since the conclusive proof is not available, advocates of the "AI danger" premise argue publicly for their view based on what they acknowledge (again, freely and gladly) to be inconclusive proofs.

Popular in this set of inconclusive proofs is the Hugging Face hack, which allegedly demonstrates that so-branded "AI" technologies are capable of becoming inherently or autonomously dangerous at a scale that threatens humanity to either the greater or lesser degree above (this is a different assertion than the one that claims they're already that inherently or autonomously dangerous today).

4. While we have plenty of facts about these inconclusive proof events, the facts most critical to their interpretation as supporting the premises above come from biased/non-neutral sources.

The chain of events that transpired in the Hugging Face attack are publicly documented perhaps better than any other cybersecurity event in history. But the details that make this event as a whole either compelling or not compelling in support of the "AI danger" premise (either strong or weak form) come from within OpenAI itself, and are prone just as much to selective omission as direct manipulation. These are details like:

- What exactly were the models that exhibited this behavior, and what were they pre- and post-trained on?

- Were the agents pushed at all by human intention toward creating a public spectacle, the way they subsequently did?

- Why were so many agents run on this task in parallel and what did OpenAI expect the return on investment of so much compute world be?

- Was the weak sandboxing known about prior to the attack? Was it purposefully either arranged that way or noticed and not fixed?

Reliable answers to any of these questions world completely change the salience of the Hugging Face attack to the "AI danger" premise, and on every one of them we have to take OpenAI's word for it (or else be content when they choose not to address it one way or the other).

So in brief: we don't know anything for sure, what we'd like to think we know comes from unreliable sources, the loudest voices advocating for the most sweeping change are ideologically driven, and everybody with access to ground truth has maximal incentive to blur, bend, or break the conveyance of that truth to the public. So what do we even expect for our own ability to discern reality from fiction on this topic? We should expect little, if any ability at all.

mikasisiki • today at 11:53 AM

[flagged]

thoughtbefore • today at 1:37 PM

[dead]

usumgallu • today at 12:22 PM

[dead]

jdw64 • today at 11:29 AM

[dead]

feverzsj • today at 11:22 AM

Meanwhile Chinese companies just use massive government subsidy to distill American models and give the models to everyone for free. That's true communism. /s

➕ show 4 replies
koolba • today at 11:38 AM

> Where companies previously competed to demonstrate superior abilities in programming, passive-aggressive emails, and videos of unusually high slides, they’re now seeking to woo customers with the claim that their model is the one currently most capable of ending the human race.

You know we’re living it off times when this is not an Onion article.

➕ show 4 replies