This looks cool and the mechanism looks plausible. I found the experience of trying to understand whether the claims here are legit to be aggravating.
First of all, the whole readme section about benchmarks appears to be Claude/Codex written. What a slog to read this.
Second, they claim success on SWE-Bench Verified, but it's only on 50 tasks, not making clear how these tasks are chosen. I know from experience that you can keep selecting sets of tasks until you get a result. Also, this was run 1x, and the lift they show actually only has a p val of 0.22.
There's a famous book "how to lie with statistics". I don't think the continual posts about benchmark results here are purposeful lies, but I think it's just so easy to fool yourself (and I've been burned, most recently building http://pellmell.ai ). Also I don't think this post is the most egregious example.
we are running them on a recurring basis, so will keep on updating that 50 number. and hence as it's currently running that section gets updated by claude only.