logoalt Hacker News

A simple fix for LLM tail latency

20 pointsby oskrimlast Friday at 5:58 AM7 commentsview on HN

Comments

nine_ktoday at 9:53 PM

Sending two identical parallel requests is the classic approach. But, logically speaking, it should also double the cost.

I would send a second request if the first request fails to return the first token within, say, 1 second. Then there's a chance the first request is stalling, which is an infrequent event.

I wonder if higher-availability tiers of LLM providers do a similar thing internally.

dvaplimatoday at 9:49 PM

Nice turn around, does anyone has a benchmark regarding other types of requests (priority vs send twice) other than voice/call? Or the tests already test that?

crisnobletoday at 10:01 PM

Why not send it thrice?

ramon156today at 9:51 PM

for a tier thats twice the cost i would expect >2x the speed. somewhere 5-10x

e.g. 1.40m would become 0.30s.

do people really pay for these priority plans?

eigenblaketoday at 9:06 PM

I love this. Simple. Useful. To the point. If AI was used, I can't tell because it is clearly representing the author's beliefs.

show 1 reply
moffkalasttoday at 10:06 PM

If you want a controllable and predictable system, host it yourself. APIs will always have outages, delays and breaking changes every so often. That's the price you pay for not doing it properly and outsourcing your job.