Is there any resource or benchmark actually comparing how different harnesses perform with a given model?
Really just feels like endless FOMO with how fast the iteration cycle is for harness and model development.
I found this: https://artificialanalysis.ai/agents/coding-agents#harness-c...
I don't think it's very good though. As an example I can use one CC instance to delegate to several to achieve complicated/open-ended goals.
That just isn't possible with other harnesses, and it's definitely not benchmarked.
I found this: https://artificialanalysis.ai/agents/coding-agents#harness-c...
I don't think it's very good though. As an example I can use one CC instance to delegate to several to achieve complicated/open-ended goals.
That just isn't possible with other harnesses, and it's definitely not benchmarked.