>Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware
That's low hanging for you?
Maybe they meant in the sense of “untapped potential”, because so far a lot of the focus has been on increasing model capabilities, not necessarily performance/power budget.
We already have examples of LLMs running 16k+ tokens a second using custom ASICs.
It's down at the moment (Not sure if it'll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens/sec. It was amazing to use, you'd no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.
I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we're currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.
We really don't always need newer better faster stronger models, there's quite a lot of room for "good enough" where getting 17kt/s at significantly lower power would be amazing.
It's likely quite close. There are so many papers proving concepts that would bring this, they just haven't been combined in production.
> That's low hanging for you
An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.
Yes. It's well within realm of possibility, but so far wasn't pursued because the Big Vendors went all-in into capability growth (rightfully testing "the bitter lesson" to its limits) and got themselves stuck in an arms race, while everyone else is barely keeping up and/or starstruck with fascination, exploring what these models can do.
This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.
That's a high hanging fruit, not low.
If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].
[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document