nice work and paper, seeing more harness benchmarks emerge and we definitely need more. I ran mouse on the Frontier Harness benchmark and scored the highest pass rate, however I'm not convinced the results there actually translate to meaning the "best" harness in practice.
Always looking for more harness evals, although I'm going broke running them across all these models and tasks.