logoalt Hacker News

mofeien • today at 12:39 PM • 6 replies • view on HN

When, in the past three years, has model progress seemed to decelerate to you, indicating some limit?

The Statement on AI Extinction Risk is more than three years old, signed by the three CEOs: https://aistatement.com/work/statement-on-ai-extinction-risk

They have been warning about AI extinction risk for years, and AI progress has only been accelerating.


Replies

nsagent • today at 1:03 PM

It's pretty telling that even with RL post-training the big labs have essentially made little progress on the hallucination rate of models. The issue is fundamental to the current paradigm, contrary to humans.

GPT-6 Astra (max) has a hallucination rate of 51% and Claude Opus 5.5 (max) has a rate of 59% according to Artificial Analysis [1].

  AA-Omniscience Hallucination Rate (lower is better) measures how often the model answers incorrectly when it should have refused or admitted to not knowing the answer. It is defined as the proportion of incorrect answers out of all non-correct responses, i.e. incorrect / (incorrect + partial answers + not attempted)
Full speed ahead like an idiot savant trying a thousand different possibilities, though half of which are without basis in reality.

[1]:https://artificialanalysis.ai/evaluations/omniscience#omnisc...

➕ show 4 replies
dml2135 • today at 2:35 PM

I actually do think it has been decelerating a bit recently. It’s just that last 1% feels much bigger than the previous 10%.

palmotea • today at 12:43 PM

> They have been warning about AI extinction risk for years, and AI progress has only been accelerating.

So, they're either liars or homicidally reckless.

➕ show 5 replies
nkrisc • today at 1:04 PM

So either they’re full of shit or we need to stop them by any means necessary.

➕ show 1 reply
sscaryterry • today at 1:27 PM

You clearly don't understand the underlying fundamentals of LLMs, harnesses, agents.

Its 100% human doing. A human set a task, a human didn't monitor it. I for one, can do jack-shit security or defensive work with Opus/Fable/Astra/Sol. Implication: Different set of rules for us, and for them. Of course running it without any checks is not going to end well, it doesn't mean its going to kill us all.

PunchyHamster • today at 1:26 PM

the ceiling of abilities seem to be growing steadily but the floor of errors seems to not change. New models can do more and more but still fail at seemingly (to human) simple tasks