How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
[flagged]
Part cherry-picking of benchmarks, part leapfrogging