except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience
You'll find that hard to prove objectively and conclusively.
You'll find that hard to prove objectively and conclusively.