I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.
It is more fun than serious at this point. Don't overthink it :)
Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point.
(Next up is the comment saying that the labs are clearly training for the benchmark.)
it's a tradition
If everyone agreed with you, the comment would disappear near the bottom of the thread
I like the benchmark. Yes, it's near saturation for SotA models, but still quite good to show where smaller models stand in relation to SotA
In this instance, I see a great image, but consistently clipping mudguards (both in 3.8 flash and 3.7 flash)