Anecdotally it's a mix. Claude is so good in part because the models are clearly trained to use the harness, and the harness (despite questionable UX) is really best-in-class when it comes to its functionality.
> run a web search, write a draft perspective from three points of view, and structure data around it
Ironically that's not harness-heavy at all, is it? Apart from sterring via system prompt, that's largely relying on the model itself to reason through the task (what to search for, which links to follow) and then synthesize the information and present it in a way that meets the user's request. Seems like a good test of pure LLM capability to me.
I find it hard to believe that if GLM 5.3 struggled with that task in, say, Pi, it would do any better in OpenCode. Unless you're talking about some next-level research stuff to provide strong guidance/steering and context offloading.