logoalt Hacker News

TomGarden • today at 6:56 PM • 3 replies • view on HN

Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?


Replies

petu • today at 7:01 PM

That's max output tokens per response limit, separate from context length

simonw • today at 7:11 PM

It's the output token limit, which has been 128,000 for Claude models for quite a while note

➕ show 1 reply