logoalt Hacker News

onlyrealcuzzotoday at 6:31 PM5 repliesview on HN

This is awesome, but tokenization is typically <0.1% of total inference time.

Presumably there's a host of applications that just need to tokenize, though, and this would be great for those!


Replies

scottchatoday at 7:22 PM

I run an AI platform and we need to tokenize fast and early to make a lot of decisions on the subsequent steps (things like routing, rate limiting and such). Its really important to do this efficiently even though its not a large % of total end to end time for the request.

show 1 reply
pipsterwotoday at 7:03 PM

1/1000 of inference compute is a non-trivial workload at scale. Gartner estimates ~$28B in inference spend for 2026 making this a $28 million dollar per year workload (edit: based on the assumption above)

Source: https://www.gartner.com/en/newsroom/press-releases/2026-07-2...

show 1 reply
noahbptoday at 8:16 PM

Time to first token, especially for smaller models, can be sharply reduced.

Latency can be just as important as overall throughput, especially for inference providers like Groq and Cerebras.

show 1 reply
GenerocUsernametoday at 7:02 PM

Always good to make it 0.001%

brcmthrowawaytoday at 7:42 PM

[dead]