We found an approach to get Jev-like properties from standard LLMs like GLM-5.3-Flash.
The core idea is to craft the input prompt so that the first output token answers the question. This makes it possible to get a decision with a single forward pass.
In the blog post, we describe the approach in detail for GLM-5.3-Flash and vLLM. We benchmark this setup against Jev and Laya. We find that our setup is on-par with Jev in terms of accuracy and speed and that it substantially outperforms Laya.
Still, in terms of costs per decision, Jev is several x better than our setup. In turn, our setup supports vision inputs.
how is Jev cheaper if I can run locally. 0.5% prefill, 0.1% decode, 99.4% cached, latency is <20ms
If you are using an autoregressive decoder (which glm is) it is not “jev-like”. You lose all of the speed advantages that Jev has.
My question is why not use Jev instead? It's faster and cheaper.
Isnt this obvious ? I would have thought people would try such things before deciding they need something like Jev
Is this a joke? “Jev-like” properties? People have been using LLMs as classifiers or rankers in a similar way for ages. I feel like we’re losing our minds
RIP Jev
[dead]
Everyone is doing this to emulate Jev, but...
I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.