Everyone is doing this to emulate Jev, but...
I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.
That's not true. I ran Cerebras as an experimental ultrafast Jev and it was faster.
was the answer correct?
i have tested jev for my use cases and its horrendously wrong, but then the follow up from jev's team is "oh, you need to boil the question down further". it's a spiral of how much do you wanna dumb down the ask so that it answers it correctly. i'll pass for now.
also, 30k input tokens is a lot.
Also, it processes all questions you ask it in parallel, which is also not possible with normal LLMs.
This has “/dev/null as a service” vibes…
I imagine it’s not so hard to optimize a model for this use case.
Off the top of my head, I would skip all the modern linear attention / state space stuff and use classical attention. But run prefill in a fully sliding-window mode so that “state” tokens simply don’t attend to far away tokens, or maybe also allow everything to attend to the first few tokens (and train like this). Now prefill is almost embarrassingly parallel, and you can make it fully parallel by duplicating work at block boundaries. (I’m not saying this is an awesome architecture if you want excellent results, but I’m also not convinced that Jev gives excellent results…)
The let queries attend to everything.
And architect the stack around this. Don’t try to cache the KV data — process the queries as you go so that the each input block and layer’s K and V data is computed, attended to, and discarded.
I’m curious whether Cerebras actually is a good device for this. Cerebras is kind of low on RAM, but if you don’t need to store KV data, maybe the entire computation fits on the die.