logoalt Hacker News

irthomasthomastoday at 5:44 PM1 replyview on HN

Why does anthropic change the set of benchmarks they use with every new model release?

https://www.anthropic.com/news/claude-opus-4-7

https://www.anthropic.com/news/claude-opus-4-6


Replies

pietztoday at 5:55 PM

1. Benchmarks saturate 2. They select the most impressive improvments