My point is that the differences between these models are so minor that obsessively benchmarking them comes across as navel-gazing.
The evidence that proves a model is actually a step function change is these benchmarks.
If a model isn’t a step function change? Welcome to research.
like all good science, measure everything
The evidence that proves a model is actually a step function change is these benchmarks.
If a model isn’t a step function change? Welcome to research.