logoalt Hacker News

himata4113yesterday at 9:10 PM1 replyview on HN

They have not increased in capabilities, they have increased in specialization.

If you train a small model in another domain it will begin losing capabilities in the former domain. This is effectively the sigmoid problem.

Although I will admit that if we discover a higher information density algorithm that it might change, but not by a substantial amount to where "super intelligence" in 1gb would be possible.


Replies

Philpaxyesterday at 9:57 PM

Over the last two years, this weight class has doubled its scores and/or saturated several benchmarks in the Qwen lineup alone without loss of generality: https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-7...

There is undoubtedly a limit somewhere (there is only so much you can pack into a given size) but it's really not particularly clear where that limit is. I don't think it's superintelligence - that much I agree with you - but I think "We already have a 1gb model that is as capable as it will ever be" is strictly false.

show 1 reply