logoalt Hacker News

LLMs are real, AI is fake

60 pointsby danaristoday at 1:47 PM25 commentsview on HN

Comments

lhltoday at 4:35 PM

I think that Doctorow, Zitron, and other "denialists" are doing a real disservice to their audiences and it's only going to make the future shock worse.

The basic claim that HF incident isn't evidence of consciousness or a spontaneous desire to hack? Sure, there was a terminal objective assigned. However, everything else beyond that strawman? Pretty shaky, IMO.

If you look at the OpenAI, METR reporting (and related collusion.wiki , rubyhack.ai reports) we are seeing strong evidence of operational agency, instrumental goal formation, spontaneous swarm formation and collaboration, capability amplification and unexpected consequences of network effects, deliberate/acknowledged violation of task boundaries. To collapse that down into "a Python loop and a chatbot" or still talk about "consulting its training data" seems dangerously shortsighted, and from my reading, demonstrably wrong from what was extracted from the logs and bot interactions.

BTW, a lot of his arguments are based on things that are factually wrong. ExploitGym has explicit instructions to only exploit target X using vulnerability Y. Everything the swarm did was by definition misaligned/against instructions.

Before his enshittification train, Doctorow used to say "don't savvy me" a lot. Hey Cory, don't savvy me. This is new emergent behavior, it's incredibly alarming and I don't think even the people paying the most attention to this field can agree or see where this is really leading to. This stuff should be in the headlines, it's unprecedented and I don't think existing mechanisms/institutions are anywhere near adequate, considering how in the dark they are responding to what's been happening.

show 1 reply
IanCaltoday at 3:36 PM

> The chatbot consults its training data

Err, no? That's not at all how llms work.

> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

They worked out how to fake the scoring, then hacked into a different system (which required finding a bunch of other exploits) in order to find the actual answers, and were trying to modify their own logs to hide what had happened.

This isn't a case of them saying "hack into X... OH NO IT HACKED INTO X".

> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

It wasn't a rival server though, was it?

> That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition.

They also tried to modify the code in the benchmark. Are you allowed to try and break into things to change the problem? edit - the agents transcripts show some of them explicitly saying that attacking HF is not allowed as part of the challenge

This all seems to dramatically underplay how interesting the actual attack was and what built up to it.

https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

show 2 replies
quicklywilliamtoday at 3:43 PM

My takeaway: We should be not be concerned about AI bots' "goals", we should be concerned about the goals of the companies making them. Powerful but not sentient technology in the hands of reckless accelerationists is a plenty dangerous enough thing.

show 1 reply
jsnelltoday at 3:39 PM

I honestly don't even understand what straw man Doctorow is arguing against here.

But he is wrong on the facts: these incidents were not merely the models already being in a infosec context and escalating beyond the intended parameters. They happened also with no kind of security elicitation. So the task was something like searching the internet for economic statistics, not to hack into a system.

reissetoday at 3:50 PM

I don't know, it reads like a pure copium at this moment.

While Zitron continuously whined about the "AI bubble", and how the models were not improving, and how spectacularly it should've blown, these same people who Doctorow accuses of, quote,

"cooking their brains by locking themselves in the bathroom, holding flashlights under their chins, and saying "Aaaaaaaaaay Eyeeeeeee" until they wet themselves in terror",

unquote, kinda promised, among other things, to give everyone an APT-level big and automated hacking bazookah, and now surprise-surprise three years later they delivered exactly on that promise, and now Doctorow is victim blaming everyone around that they were unprepared for that!

But it was you who said it was all hype, smoke and mirrors, it was you who said the AI is quote,

"a product of limited utility that has been shoehorned into high-stakes applications that it is unsuited to perform",

unquote, why are you suddenly surprised everyone around was not inspired to do anything around it?

The real world software threat model was never suited for a relentless hacker-ex-machina, limited only by the token count you can throw at the task. And maybe we didn't prepare in time because no one believed in possibility of such a machine, because people like you said it was just for-profit scaremongering?

Give or take, AI evangelists gave us all a pretty wild and unbelievable set of expectations few years back. I didn't believe them back then too. But now they're steadily delivering on _some_ of them, and we should be correcting our world model to take into account _all_ of them might be true, instead of making up reasons why other predictions will certainly fail.

show 1 reply
daishi55today at 3:40 PM

> Once you understand the corporate culture of AI "hyperscalers" consists primarily of everyone cooking their brains by locking themselves in the bathroom, holding flashlights under their chins, and saying "Aaaaaaaaaay Eyeeeeeee" until they wet themselves in terror, a lot of things snap into focus:

This is the writing of someone who has absolutely zero interest in or intellectual curiosity about the subject of their writing.

sigmartoday at 3:08 PM

>To understand the truth about the Hugging Face hack, you could do a lot worse than to listen to Ed Zitron and Cal Newport's recent podcast conversation

Lol, okay...

The crux of this piece is Doctorow saying that the hack was just a stochastic parrot repeating steps it has been trained on. Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened. "Oh, it only made those paperclips because it saw instructions on making paper clips in the training data." These are some 2024 arguments...

show 5 replies
w22ooptoday at 3:40 PM

Hm. every firm with Atom enginering is in this same situation. Working on project, and this project can destroy earth

danaristoday at 1:53 PM

Some really important and informative stuff in here—I certainly had no idea just what the nature of the prompts and tooling that produced the HuggingFace exploit were.

This shows fairly clearly that (as I already suspected) this was not, remotely, an LLM "going rogue." This was humans planning poorly, not thinking of the consequences of their actions, and giving LLMs too much scope and a lousy prompt.

show 3 replies
aaron695today at 3:44 PM

[dead]