Putting only the tasks and results in the repo is a poor decision. These conclusions would be far more credible if anyone could re-run the benchmark.