logoalt Hacker News

elahiehtoday at 12:51 AM0 repliesview on HN

One reason might be that Claude Opus 4.7 thinking benchmarks better on Arena Coding at https://arena.ai/leaderboard/text/coding ... hopefully that effectively assesses correctness. It doesn't account for reliability though.