logoalt Hacker News

seizethecheesetoday at 6:01 PM1 replyview on HN

Wow, this is really weird. I looked at the github and see you published benchmark updates in these sizes of N:

9 → 20 → 36 → 50

These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:

> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.

> You chose different sized task groups (almost certainly not the case).

> You excluded timed out trials, or something (also would be dishonest).

Hopefully there's a better explanation here!


Replies

shrishdwitoday at 6:10 PM

that's just when our claude limit's were about to exhaust :/ as we have to spin up a whole claude session to test it.

but you could 100% be Sherlock Holmes!

show 1 reply