I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.