Would creating new benchmarks every month solve this problem?

operatingthetan • yesterday at 6:28 PM • 1 reply • view on HN

Replies

Or create "blind" benchmarks.

10 groups of 3 researchers, all have their own benchmarks that they do not share (testing it without the authors knowing is a different problem, maybe they only run the benchmarks when the gen-pop has access to the models).

that's 10 different tests. Aggregate pass rates

alt Hacker News

Replies