logoalt Hacker News

markasoftwareyesterday at 7:54 PM2 repliesview on HN

It's well known 35b is much faster (on any hardware) and quite a bit dumber


Replies

dofmyesterday at 9:17 PM

This really very much depends on how you are using it, I think. If you intend to leave it to solve long context problems and write whole prototypes, the 27B is going to be much better.

But if you are sort of pair-programming with the model, the speed obviously matters and I think then the 35B is acceptably smart, and when it's wrong it'll be wrong much more quickly. It seems very good on SQL and PHP, and I assume on typical JS and Python.

I would rather work that way, so I hope they do produce a small MoE model.

LoganDarktoday at 11:16 AM

I had no idea. Where can we learn stuff like this?