logoalt Hacker News

yipinwong • today at 5:07 PM • 8 replies • view on HN

A question someone not trainined in AI/ML field, Is a decision model that easy to crete that there are floods of these JEV alternatives already?

Or are companies/people already building this based on say an arXiv docs? n

---

The pricing is ... hm more expensive but not at the point I won't give it a try due to the embeded vision encoding


Replies

TeMPOraL • today at 5:37 PM

Yes, it's easy. The thing people are missing (especially those believing AI is a "dead end" and "not transformative") is that the field has been advancing so fast in the past few years, that there's lots of such unexplored avenues, unpicked low-hanging fruits, that everyone just raced past. We've barely begun exploring the capabilities ML brought us - patterns, applications, and architectures.

Now that we're hitting against the hardware supply limits of global economy, I expect more people to go back and revisit the things left along the way in the mad rush to "just throw more compute at it / make a bigger model" - and thus many more cases like Jev to show up in the next few years.

nico • today at 5:31 PM

The basics are pretty simple. And depending on what your specific need is, the model can be really really basic, fast and super effective (ie. run on a mobile device and process thousands of requests in <100ms)

I've been playing with this for the last year or so. Started with a personal email classifier, also did benchmarks with some public datasets, then created a couple classifiers that could play Doom, and now I've been trying out some other experiments, like a request proxy/router to automatically choose a classifier and fallback to LLM to handle unseen requests

Jev did a great job at creating hype, but also at shaping the concept and space of "decision engine" or "decision model". People were already doing this with LLMs, which is very inefficient for most tasks like that, and the Jev guys figured there was a market there. It seems like they were right, and now there's a rush to flood the space, taking advantage of the hype window

calebkaiser • today at 5:35 PM

There is a bunch of stuff to tease apart.

In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.

A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.

I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.

But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.

The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

➕ show 1 reply
conmod278 • today at 5:30 PM

Live coding Jev from Scratch | Understanding Qwen architecture

https://www.youtube.com/watch?v=AzxoU7kxjig

orbital-decay • today at 5:50 PM

Yes it's easy for an established shop, all they need to do is to tweak the post-training workflow. "Decision model" is the same kind of marketing as "LRM" attempted by OpenAI when RL CoT was new (to hyped up crowd). It's still fundamentally a classifier used for "decision making", games and RP were using generalist models and constrained outputs to do what the DOOM demo does for years.

janalsncm • today at 6:06 PM

The interesting part is also the easy part. The model and architecture are not hard for an experienced machine learning engineer to build.

The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).

XCSme • today at 5:15 PM

You can make a basic one in minutes based on existing open-source models.

Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.

Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.

It's not really a new technology, it's more like a new use-case.

➕ show 1 reply
redox99 • today at 5:23 PM

Yes it's very easy if you have fairly basic ML knowledge.