logoalt Hacker News

seaalyesterday at 5:30 AM1 replyview on HN

Is there any resource or benchmark actually comparing how different harnesses perform with a given model?

Really just feels like endless FOMO with how fast the iteration cycle is for harness and model development.


Replies

shepherdjerredyesterday at 5:58 AM

I found this: https://artificialanalysis.ai/agents/coding-agents#harness-c...

I don't think it's very good though. As an example I can use one CC instance to delegate to several to achieve complicated/open-ended goals.

That just isn't possible with other harnesses, and it's definitely not benchmarked.