In my case it actually made a difference. I tested it on my own code review set with Gemini Flash. On low it found fewer bugs than on default around 90% versus 97% but was about three times faster and a lot cheaper. On high it actually found everything in the hardest case but took almost three minutes per call. so I think the difference is real you only see it if you run the same fixed cases a few times and not by feel. And yes I used Gemini rather than the GPT Sol that is mentioned here would be interesting to see the same tests on Sol.