Wow, this is really weird. I looked at the github and see you published benchmark updates in these sizes of N:
9 → 20 → 36 → 50
These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:
> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.
> You chose different sized task groups (almost certainly not the case).
> You excluded timed out trials, or something (also would be dishonest).
Hopefully there's a better explanation here!
that's just when our claude limit's were about to exhaust :/ as we have to spin up a whole claude session to test it.
but you could 100% be Sherlock Holmes!