logoalt Hacker News

bobbylarrybobbytoday at 4:42 AM2 repliesview on HN

The models themselves have far from plateaued. Maybe someone finds a way to get a really capable model down to, say, 12GB of ram. Then we'd be in business.


Replies

swiftcodertoday at 6:20 AM

Agreed. We've just seen DeepSeek post-train their ~300 billion parameter flash model to outperform their 1.6 trillion parameter pro model, in the space of a few months. There would seem to still be quite a few opportunities on the table to bring big model smarts down to the smaller models