logoalt Hacker News

wonnagetoday at 6:20 AM0 repliesview on HN

> Overall, the best-performing model was Claude Opus 5 on “reasoning” mode, which still made mistakes in 39 per cent of answers.