It seems that the agentic search benchmarks have fallen victim to reward hacking, just like the coding benchmarks. It's good to see mitigation efforts being made to address this problem.
What do you foresee in the future releases and improvements to this?