logoalt Hacker News

amlutotoday at 1:28 AM0 repliesview on HN

> And yes, the paper is already outdated

The paper is about the fact that the benchmark’s evaluator is prone to egregious incorrect rejections of what answers that it should accept. A new model will not invalidate that issue.