logoalt Hacker News

alyssamaruyesterday at 4:52 PM1 replyview on HN

How much does the harness really matter for evals?


Replies

wittydeveloperyesterday at 5:00 PM

A lot, especially for performance. That's why we built our own benchmarks that account for both the model and the harness. For example, with Claude Opus 5, the accuracy gap can be up to 3% and performance up to 200ms, depending on whether you're using Deep Agents, Eve, or Fx.

More details here: https://www.stagehand.dev/evals