logoalt Hacker News

seizethecheese • yesterday at 5:35 PM • 7 replies • view on HN

Name a few of these low hanging fruit left around.


Replies

TeMPOraL • yesterday at 5:48 PM

Jev is one.

Diffusion transformers are not "easy" but underfunded.

Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

Or imagine automated sliding doors that don't suck.

--

[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.

➕ show 7 replies
hobofan • yesterday at 5:54 PM

Closely connected to decision models: A good library to do ranking based on pairwise ranking on multiple attributes. By using a decision model (especially one that can make decisions on multiple fields at the same time) this becomes a lot faster and more powerful. Could make for a pretty nice search reranker as well as prioritizer for many problems.

Of course you can also do ranking one-off with a decision model, but this likely less stable, and by doing pairwise ranking you can also relatively quickly do incremental inserts to the list.

sarkarghya • yesterday at 7:10 PM

I can imagine advancements on making smaller models work together better instead of a generalized core. Imagine a community or city having a https://pirateface.co/ so that the shard of the model that you need can be streamed in with minimal latency with your box only holding the minimal version (say deepseek v4 flash as orchestrator) of the model that you use on day to day basis.

We have overcome split brain problems before so this wont be our first

➕ show 1 reply
murkt • yesterday at 5:37 PM

Easy to reach doesn’t automatically mean “easy to see”.

sroussey • yesterday at 9:53 PM

Just look at all the model type on hugging face. LLMs are a small percentage.

CamperBob2 • yesterday at 7:43 PM

An example I like to use is: compare the quality and scope of games released with a brand-new console to the ones released for that console towards the end of its life, when everyone has learned how to take advantage of whatever weird, wacky hardware Sony invented for that console generation.

There is still a lot we don't know about how to get the most out of existing LLM components from a speed or cognitive-performance perspective. People could easily spend the next decade studying and refining what's been built so far, even if no new, original approaches ever arrive.