logoalt Hacker News

brokencodeyesterday at 10:27 PM1 replyview on HN

Maybe in terms of code produced, but one token is only a fragment of a thought for an LLM.

It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.


Replies

titoyesterday at 10:40 PM

Oh, is that how it works? So, when somebody says a model is running at X tokens per second, it means that the thinking process is running at that, and output tokens are much lower then? Thanks to the explanation.

show 1 reply