logoalt Hacker News

audunwtoday at 1:50 PM5 repliesview on HN

It doesn’t seem like AIs are accelerating at all if you ask me. We seem to be plateauing. Smaller and open weight models are catching up to the closed weight frontier models on benchmarks. If AI labs were able to use the smartest models to accelerate their development, the likes of OpenAI and Anthropic would be accelerating away from their competitors. But no such thing is happening. The focus has shifted from intelligence to cost and speed.

The breakthroughs in mathematics are impressive. But it doesn’t feel that different from what machine learning has done with Chess, Go and protein folding. They’re finding patterns in our systems and in nature. That’s what they’ve always been good at.


Replies

perrygeotoday at 3:19 PM

Mostly I'm seeing breakthroughs in the ways we use LLMs. Agentic harnesses, MCPs, etc - its the wild west still but we've come a long way from a basic chatbot. Gains are now coming from tools that make better use of the LLMs existing capabilities, and put guardrails on their worst tendencies.

I personally feel like we've plateaued in raw model intelligence (even regressed, I find sonnet 4.6 to perform better than Opus 5) but we've gifted them new skills that allow them to run for longer and explore the search space more thoroughly, making them more effective at the same level of "intelligence".

Take two smart people and a problem to solve. Give one person the tools, the other person nothing. The one with the best tools wins. At some point its more about abilities than raw intelligence. Watching a "frontier" model fumble with basic syntax is still common, but not if you give it treesitter.

anon373839today at 2:08 PM

Agree. We’re going to hear a lot of buzz about recursive self-improvement in the near future, which I’d cynically say is meant to address the naked emperor you just called out.

Can labs get a boost augmenting more of their processes with automation? Sure. Will it result in a self-sustaining takeoff to infinity? … No. It will saturate too. Just my opinion.

skybriantoday at 2:34 PM

The math and cybersecurity improvements this year don’t look like plateauing to me. They’re clearly improving?

But I think they’re becoming more specialized. Luna is working fine for me for ordinary web development, but I was impressed by Sol tracking down an OS-level bug that was causing my Playwright tests to be flaky. Previous models couldn’t figure it out.

A lot of the improvements might be on questions that ordinary users aren’t normally asking.

spicyusernametoday at 2:11 PM

Nothing wrong with a conservative take, especially given all the hype, but I often wonder if its hard to see the magnitude of the change at "day-to-day" speed.

And the emotions one gets from reading the "takes" on news sites, blogs, etc leaves a bad taste in ones mouth and in hoping there is no change, we make it even harder to see.

show 1 reply
RunSettoday at 2:06 PM

Calling an "LLM" "AI" does a disservice to AI before it even exists.