Funny enough, I have a simple test that I have been running on each new model that catches my eye on OpenRouter. It is just a short prompt that asks the model to research yesterdays news and summarize it in a specific format along with a critique of the article or a highlight of any bias: https://gist.github.com/james2doyle/6afb04ea6b18e1a36bc45258...
I have been doing this test for about a year now. I run it on different harnesses and apps as well just to give me an idea of what they differences between them might be. Since I have been running it for a while now, I have a good sense of the correlation between this output and how the model will be on the rest of the things I want it to do.
I would say that almost every model (just tested Mercury 2.5 Preview and IBM Granite 4.2-8b) are pretty much the same on this task. Some are more diligent with how many sources they go out and get, but for the most part they are all very close in quality and will follow the instructions very well.
So saying "everything struggles" with that has just simply not been my experience.