> All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output
I keep seeing comments added to code, which reads like reasoning output instead of meaningful words. I see this behavior for both OpenAI and Anthropic models (for several harnesses as well).
But this is a sample of one. And I may be in a situation where I'm more negative to the output from LLMs in general.