As an MLE who has been failing to get anyone interested in classifiers for many years, the hype around Jev makes me scream internally.
Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing.
on the bright side, maybe this zero-shot classifier wave could be the back bone for more specialized classifiers (with more mindful selection of data and training)
I want to share the rage. Can you expound on what makes you scream?
Jev is classification for normies. Look ma, no training.
> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now
I'm no prompt engineer, but I've found them to be too slow any time I wanted to use them that way. I have never tried Jev, but apparently it is supposed to be fast, so it seems like, according to the marketing, it could become usable where LLMs haven't been.
Maybe instead of screaming you can take this as a chance to level up your engineering.
Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ.
Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from)
Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that.
> As an MLE
Yea but those models you were building were lame and inaccessible to play with for common devs.
Just because they have the same api doesnt mean you were building the same thing.
I think part of what made Jev catch on is that the API is like an if(...) or switch statement
People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal
But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank