logoalt Hacker News

anilgulechatoday at 3:22 PM0 repliesview on HN

> Qwen2.5-Coder-3B

Basing it's findings of LLM as judge on this model, and then proceeding to ignore it. This article can be safely ignored as well.

LLM as judge in harness evals is the way to go, for any of your custom needs. Design the eval well.