I completely agree. First, there's the split between the ones using AI to write coder and there's the Ai that gets used in software.
Writing code is _hard_ because it's by definition accurate logic math will tons of abstractions that build on each other. Our brains are great at this. Frontier models with lotsa context are too. Let's ignore those though.
There are a lot of problems that are simple. Accuracy is important (e.g. is it nudity?), but it's still simple. The difference between Jev and the general idea of small models you described is that Jev (and a million copycats, Jev being a copycat itself) only answers multiple choice. That's clever; why wax poetic when you can ask a question and just get back the multiple choice, standardized tests love them for a reason.
So really it's: a) "Big Know it All" Models. Slow and pricey, but they'll handle anything you through at them. b) "Small Domain-Specific" model that still speak answers. They need some amount of "can form eloquent answers" in there with the specialized knowledge so still have a bit of a floor, but totally make sense. c) "Memorized multiple choice tests, and can read but not write." Sometimes open-ended answers really are a downside when the answer space is sufficiently constrained, and read-side literacy really
Not nit-picking, but (c) really is more different than just a specialized model. And you're right, lots of problems fit that shape — hell, before AI blue up the world that's _all_ computers could do — and many times ambiguity and hallucinations are drawbacks of open-ended answers. You just need to ensure that "N/A" is always a valid answer, or at least that you get a confidence score.