logoalt Hacker News

karmasimidatoday at 7:56 PM4 repliesview on HN

Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5

Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.


Replies

_superposition_today at 8:03 PM

I must be on the wrong X/Twitter then.

show 1 reply
torginustoday at 8:37 PM

It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.

nsingh2today at 8:00 PM

Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.

show 1 reply
emp17344today at 8:37 PM

Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.

show 1 reply