logoalt Hacker News

losvedir • today at 9:46 PM • 0 replies • view on HN

> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens

Can someone help me understand this? I might have an out of date mental model of how these things work.

Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.

But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?