I am increasingly hesitant to use non-native harnesses - model providers are now starting to train their agents for use within the harness. An eval like terminal bench can only capture so much data. I don't want to have to assess each harness every model release to make sure it's working as well as it can.
I’m the opposite. I want one open source harness to rule them all
Cost efficiency is a plus
Interestingly, MiMo was trained across multiple harnesses and it improved capabilities
This is the very reason why I avoid to use native harness. They are optimized for the economic benefit of the provider, not for mine. Love Pi because it does a good job managing the context, open code meh, claude code nope.
This matters less as models get better and everyone settles on the same overall harness architectures. The model matters more than the harness anyway.
The bigger issue is that the use cases and harnesses for models is infinite, which is hard to compress into benchmark numbers that actually apply to you.
Everyone is benchmaxxing, desperate to sell, and almost nobody except the labs is doing actual science on the results, so harnesses tend to be chosen on voodoo and hunches, like which company made it. There isn't necessarily a good alternative though, bearing the cost of being a harness researcher is probably not many people's goal.
And yet, open harnesses works better and use less tokens than native harnesses trained to consume as much tokens as possible.
I mostly use GLM. But I will not touch ZAI's harness with a 10 mile long pole no matter how efficiently they couple it with GLM.