logoalt Hacker News

GodelNumberingtoday at 8:36 PM1 replyview on HN

The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:

Terminal-Bench 4.0: High (57.9%), Max (56.7%)

DeepSWE: High (73.3%), Max (71.5%)

It _loses_ 1-2% performance going to High from Max


Replies

XCSmetoday at 8:54 PM

That's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.

show 1 reply