Hasn't it been repeatedly shown that many small models perform worse than a large model of the same total parameters?