This is useful for classification problems; any time you need to write software that looks at some fuzzy data and needs to make a probabilistic decision. It's far more cost-efficient and performant to use this type of model instead of an LLM.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.
I understand but you can already do that with an LLM, it just comes down to your prompting. It is not exactly the same, but I was asking for a specific use case where that really matters