One of the authors here. We have been seeing lots of benchmaxxing and leakage in standard web search benchmarks such as BrowseComp.
We developed this live benchmark with daily/hourly sampled fresh queries matching real agentic search traffic to estimate actual search performance of different AI search providers.
It seems that the agentic search benchmarks have fallen victim to reward hacking, just like the coding benchmarks. It's good to see mitigation efforts being made to address this problem.
What do you foresee in the future releases and improvements to this?