logoalt Hacker News

Zetaphoryesterday at 8:22 PM1 replyview on HN

Open models keep getting bigger, but smaller open models also keep getting smarter. What I can do with an 8b used to require a 32b.

Also the decision makers who are signing off on things like ChatGPT Enterprise are at least 18 months behind the curve of what you can actually do with these things and how cheap they can be. They're still trying to figure out how to actually adopt the tech out of a sense of fomo, nevermind making nuanced decisions about hosting an open weights model. I see this firsthand in my own job.

I'm talking about adding facts to a model by modifying engrams or trying to bolster guardrails with J-washing, meanwhile they're still trying to figure out how to best prompt Copilot.

Give it a few years for everyone else to catch up, I'm barely able to catch my breath before there's some new development in the open source/weights space


Replies

latchkeyyesterday at 8:36 PM

> Open models keep getting bigger, but smaller open models also keep getting smarter.

Not only bigger, but smarter and more capable. From what I can tell, smaller are only getting smarter in very specific areas. There is a subtle difference there, that is extremely important.

> What I can do with an 8b used to require a 32b.

What exactly do you do with an 8b? I usually ask this question and either get no response or it is something that doesn't generate anything of value. So, please surprise me.

show 1 reply