logoalt Hacker News

kzrdudetoday at 7:46 PM1 replyview on HN

I think that's called test-time scaling i.e using more tokens at infer time to squeeze out higher model performance. That's must be part of the explanation for good benchmark results.


Replies

Casteiltoday at 8:26 PM

Yep.. it's pretty obnoxious for real-world use with the default 'xhigh' thinking. Ridiculous amount of "Wait, actually.." which might help for complex coding tasks but makes it unbearable for general purpose use.