Maybe in terms of code produced, but one token is only a fragment of a thought for an LLM.
It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.
Oh, is that how it works? So, when somebody says a model is running at X tokens per second, it means that the thinking process is running at that, and output tokens are much lower then? Thanks to the explanation.
Oh, is that how it works? So, when somebody says a model is running at X tokens per second, it means that the thinking process is running at that, and output tokens are much lower then? Thanks to the explanation.