logoalt Hacker News

seizethecheesetoday at 5:19 PM1 replyview on HN

This looks cool and the mechanism looks plausible. I found the experience of trying to understand whether the claims here are legit to be aggravating.

First of all, the whole readme section about benchmarks appears to be Claude/Codex written. What a slog to read this.

Second, they claim success on SWE-Bench Verified, but it's only on 50 tasks, not making clear how these tasks are chosen. I know from experience that you can keep selecting sets of tasks until you get a result. Also, this was run 1x, and the lift they show actually only has a p val of 0.22.

There's a famous book "how to lie with statistics". I don't think the continual posts about benchmark results here are purposeful lies, but I think it's just so easy to fool yourself (and I've been burned, most recently building http://pellmell.ai ). Also I don't think this post is the most egregious example.


Replies

shrishdwitoday at 5:38 PM

we are running them on a recurring basis, so will keep on updating that 50 number. and hence as it's currently running that section gets updated by claude only.

show 1 reply