I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do.
Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
When comparing OpenAI and Claude thats pretty much true, but not Gemini... And have you tried Antigravity? Yikes