Same. Mistral 7b has been more than I ever needed for text for years now.
Unless you must 1-shot with no harness it’s the same amount of power, maybe more because the big “good” models make too many assumptions and tend to become rigid.
Mistral 7b can do anything, and it’s basically instant even on an M3
What kind of work are you doing? For example, if I have some code in the hot path and I want to do all the usual tricks to help the compiler vectorize it, such a small model is not able to do much.
What are some limitations you have found with using a smaller model like that?
Sounds interesting. Can you give more details on your workflow and what tasks you use it for?