logoalt Hacker News

torginusyesterday at 8:37 PM0 repliesview on HN

It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.