> A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?
Wait a second, are we taking into account the massive difference in terms of resources of these two companies?
Irrelevant when they say they're competitive with Fable and Astra. They don't get to then roll that back and then say "but we have less compute!"
You're either competitive or not.