When it comes to quality of outcome, since at least Feburary, the harness has almost equal, if not more weight than the model itself. It's no longer "which model is the best?" it's "which model + harness is the best?"
I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models
It's worth nothing that recent Claude models seem to have gotten worse at tool calling outside of Claude Code and the SDK: https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/
I just wanted to emphasize this. Harness is a big part of how things perform thus usually it's harness + model co-design that's important.
Which works better for you?
> the harness has almost equal, if not more weight than the model itself
This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness.