logoalt Hacker News

fennecfoxy • yesterday at 10:26 AM • 2 replies • view on HN

Idk why Jev was even hyped in the first place. Mostly by social media people who are the type to write inspirational LinkedIn articles, Medium posts and make longwinded Youtube videos about subjects they don't know or care to know much about.

It's always been obvious that small models with a specific purpose will outplay the larger more general models. It becomes a question of "do I want to be able to toss _anything_ at this frontier model, have it handle it all but pay the price" or "do I want to toss _specific things_ at this micro model, have it handle it but not be flexible".

It's the same with agentic RAG/search, things are moving towards much smaller search-specific models rather than tossing RAG chunks at a large frontier LLM. It's like how I can ask a multimodal frontier model to identify bounding boxes for objects in an image...but if I want to do that faster/cheaper/at 60fps then I should be using yolo or similar.


Replies

schrodinger • yesterday at 6:42 PM

I completely agree. First, there's the split between the ones using AI to write coder and there's the Ai that gets used in software.

Writing code is _hard_ because it's by definition accurate logic math will tons of abstractions that build on each other. Our brains are great at this. Frontier models with lotsa context are too. Let's ignore those though.

There are a lot of problems that are simple. Accuracy is important (e.g. is it nudity?), but it's still simple. The difference between Jev and the general idea of small models you described is that Jev (and a million copycats, Jev being a copycat itself) only answers multiple choice. That's clever; why wax poetic when you can ask a question and just get back the multiple choice, standardized tests love them for a reason.

So really it's: a) "Big Know it All" Models. Slow and pricey, but they'll handle anything you through at them. b) "Small Domain-Specific" model that still speak answers. They need some amount of "can form eloquent answers" in there with the specialized knowledge so still have a bit of a floor, but totally make sense. c) "Memorized multiple choice tests, and can read but not write." Sometimes open-ended answers really are a downside when the answer space is sufficiently constrained, and read-side literacy really

Not nit-picking, but (c) really is more different than just a specialized model. And you're right, lots of problems fit that shape — hell, before AI blue up the world that's _all_ computers could do — and many times ambiguity and hallucinations are drawbacks of open-ended answers. You just need to ensure that "N/A" is always a valid answer, or at least that you get a confidence score.

codessta • yesterday at 3:50 PM

I think that technical people tend to underestimate the importance of convenience. If it were up to hacker news, everyone would be on Linux, everyone would assemble their own laptop and nobody would own an iPhone.

Even for technical problems sometimes having something fast and convenient that doesn't require a ton of fine tuning opens up a lot of doors. Those door might've easily been openable previously with a few days of effort. But the difference between a few days and something that you can set up with a quick account and a prompt is massive