What harness do they use for testing?
Back in 2025 it was common to test models in a different harnesses.
I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc.
Why did it come out of fashion ?
EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%
Details about the methods can be found here: https://artificialanalysis.ai/methodology/intelligence-bench...
Specifically they use this harness: https://github.com/ArtificialAnalysis/Stirrup