That does not match my own experience, which is why I wonder if Google has evidence of that.
Consistently, lower intelligence models provide worse results in my own work. But I don't have evals on my side, just vibes.