Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5
Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.
Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.
Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.
I must be on the wrong X/Twitter then.