I'm not sure how you could run such a benchmark without leaving it possible for the labs to easily detect and fudge the results.