logoalt Hacker News

drbscltoday at 5:29 PM2 repliesview on HN

Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort


Replies

persedestoday at 8:19 PM

so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)

93potoday at 6:21 PM

Is verboseness the only measure of token efficiency towards overall task completion?