logoalt Hacker News

strandstantoday at 4:31 PM1 replyview on HN

The scores are in blog post's bar chart. For Terminal Bench 2.1, Strands harness (Fable 5) scored 69.7 while Claude Code (Fable 5) scored 61.8. This is on high effort.

I hear you tho about saturation. We're working on a follow-up deep dive post with more harnesses, so could look into Terminal Bench 4.0?


Replies

seizethecheesetoday at 4:40 PM

Then I'm really confused. Terminal Bench 2.1 scores on Artificial Analysis are like 80-90%. https://artificialanalysis.ai/evaluations/terminalbench-2-1