logoalt Hacker News

LoveMistraltoday at 5:12 PM3 repliesview on HN

Same. Mistral 7b has been more than I ever needed for text for years now.

Unless you must 1-shot with no harness it’s the same amount of power, maybe more because the big “good” models make too many assumptions and tend to become rigid.

Mistral 7b can do anything, and it’s basically instant even on an M3


Replies

frigidwalnuttoday at 5:19 PM

Sounds interesting. Can you give more details on your workflow and what tasks you use it for?

show 1 reply
Almondsetattoday at 5:29 PM

What kind of work are you doing? For example, if I have some code in the hot path and I want to do all the usual tricks to help the compiler vectorize it, such a small model is not able to do much.

show 1 reply
casper14today at 6:18 PM

What are some limitations you have found with using a smaller model like that?