One thing I've noticed and HATE, is that when you increase thinking-effort, that seemingly increases response-length. Meaning that X.High is longer than High, which is longer than Medium, etc.
Which is kind of the inverse of how people work; a really smart person can condense difficult ideas into simple[r] terms. Whereas people who struggle speak a lot but say very little.
High/X.High do seem to deliver better quality results, but it sometimes feels like needle-in-haystack extracting that from the word vomit.
I just go over the comments with Gemini 3.1 Pro at the end which has a much more normal "voice" and it doesn't lose nuance as a cheap model would. I don't care so much about what Claude writes during the debugging as I just do all the cleanup at the end instead of at every commit.
It's so bad I've made myself a Pi extension that rewrites responses in side by side view using models on Cerebras (insanely fast tps)
The higher the effort the more things Claude checks, and it's eager to tell you about all of them
See, this insight it had early on looked like a red hering for a while, but then turned out to be load-bearing. And that's not just a difference in semantics, it changed the whole conclusion (spoiler: it didn't). And Claude is very eager to tell you about this exciting journey
OpenAI has separate dials for verbosity and reasoning_effort (but could still do a better job).
I hate this too, I had to switch to Codex, because the skill to force Claude Code not to think too much about very, very basic things no longer worked
With LLMs, you're still mostly read things "off the tip of the tongue". A better comparison is observing a smart person talking to themselves while working on a tough problem.
EDIT: also there's a reason the dial is called "effort", not "smarts".