logoalt Hacker News

lambdayesterday at 6:59 PM3 repliesview on HN

> I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now

10-15 years? The current rate is closer to 10-15 months.

15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

Today, you can easily run Qwen 3.6 27B on a variety of consumer hardware. It scores 37 on that index.

Here are a number of open weights models that you can run locally compared with the frontier class models from 7 to 15 months ago: https://artificialanalysis.ai/?models=o3%2Co3-pro%2Cclaude-4...

I've run all of these models on my laptop (Strix Halo, 128 GiB of unified RAM); the bigger ones, like MiniMax M2.7 and DeepSeek V4 Flash, need to be done at fairly aggressive quants that will certainly lose some performance and not quite hit the performance of the unquantized models. But still, it's definitely the case that you can run models that are competitive with the frontier models of 10-15 months ago on consumer laptops.

Heck, just announced though the weights haven't yet been released for independent confirmation is MiniCPM5-2B, a 2 billion parameter (small enough to run on your phone) model, that according to their benchmarks has performance competitive with GPT-4o, a frontier class model from 2024.

https://nitter.net/i/status/2079088670804767114

So that's around 1 year for frontier to consumer device class, 2 years from frontier to phone.

Now, this kind of rate won't necessarily keep up; it's possible that local models will hit a performance ceiling before frontier models do. There's only so much information you can cram into a certain number of bytes, and the AI boom is causing hardware prices to skyrocket so keeping consumer hardware from advancing quite as fast as it had been.


Replies

345f506fdc718yesterday at 7:05 PM

> 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors

show 1 reply
holodukeyesterday at 7:06 PM

No. We need objectively around 192 to 512gb of very fast memory to be able to run really useful models. I don't see local hardware with these specs coming in 1 to 2 years. There are a big number of initiatives currently taking place to increase ram output. But it will take another 3 years minimum to close the current supply issues. China is fast pacing forward to have its own chip baking factories with small enough nano scales to have fast chips. Will also take a few years.

Forgeties79yesterday at 7:01 PM

> 10-15 years? The current rate is closer to 10-15 months.

The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.

show 2 replies