logoalt Hacker News

The efficient frontier of LLM inference

58 pointsby philipkielyyesterday at 11:48 PM9 commentsview on HN

Comments

ttoinoutoday at 1:08 AM

   Inference techniques either move a deployment along the latency–throughput frontier or push the entire frontier out, creating more efficiency to allocate.
This is a tautology. You can say that with anything. Gastronomy techniques will make a previous recipe better, or create a new recipe better than others, or a mix of both.
brrrrrmtoday at 12:23 AM

this is a nice and concise writeup. what's striking to me is that these techniques really have not changed in /years/. sure, precision has become slightly lower, spec decoding acceptance has gotten slightly better and the complexity of parallelism is trickier with mixture of experts. but no new concepts in a very long time!

the absolute most impactful improvements for inference comes at architecture design time. I firmly believe everyone who cares about impacting model efficiency should look there

show 2 replies
datadrivenangeltoday at 12:57 AM

The author does not deeply mention that quality/intelligence is a third dimension here in addition to throughput and latency, and the frontier is jagged so quality and intelligence require bespoke benchmarks to evaluate tradeoffs for speed and cost.

show 1 reply
calclaviatoday at 1:04 AM

good recap on the recent inference techniques!

paidxtoday at 1:07 AM

[flagged]

jing09928today at 1:33 AM

[dead]

nedo_vartoday at 12:40 AM

[dead]

killerdog10today at 2:03 AM

[dead]