logoalt Hacker News

travishcronintoday at 8:53 PM0 repliesview on HN

Interesting to see harness benchmarks for coding agents. The same problem exists for conversational and data agents but I dont see anyone benchmarking them yet. Seems like manual-spot checks are the norm