we are running them on a recurring basis, so will keep on updating that 50 number. and hence as it's currently running that section gets updated by claude only.
Wow, this is really weird. I looked at the github and see you published benchmark updates in these sizes of N:
9 → 20 → 36 → 50
These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:
> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.
> You chose different sized task groups (almost certainly not the case).
> You excluded timed out trials, or something (also would be dishonest).
Hopefully there's a better explanation here!
Wow, this is really weird. I looked at the github and see you published benchmark updates in these sizes of N:
9 → 20 → 36 → 50
These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:
> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.
> You chose different sized task groups (almost certainly not the case).
> You excluded timed out trials, or something (also would be dishonest).
Hopefully there's a better explanation here!