Ah, yes -- A closed source benchmark that Anthropic paid for that Anthropic ranked highest.
0/10
Just by seeing the title I knew it was going to be like the meme of Obama giving himself a medal. Why are they even doing this? Are there still people that trust LLM benchmarks made by the LLM companies? It seems to me that the only people still using Claude are the ones that have it for free at work (me). Some of my colleagues even started using their own private OpenAI subscriptions to avoid using Claude, others are using Gemini flash to decipher what Claude is saying...
Yep, conflict of interest is the elephant in the room and its absence from the "reducing risks from advanced AI" point list is conspicuous:
* Figuring out how to prevent the incentives of frontier AI labs from aligning with anti-social deployment of AI rather than pro-social deployment of AI
It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. They put in place organizational structures to control it and then promptly smashed the structures once they smelled money. They already failed the integrity check. Even if their AI told them what they didn't want to hear I'm sure they would ignore it. Maybe they already have.
The conflict of interest is real there, but some benchmarks really ought to be closed source. Otherwise, the second your benchmark is public labs will overfit their new models on it and it will cease to be useful.