logoalt Hacker News

JSR_FDEDtoday at 3:02 AM0 repliesview on HN

I like the 2x2 grid that describes when to fine-tune a model, when to use a frontier model, etc.

From the article it’s not clear how the scorer grades every episode - was it a frontier model that assigned the grade? How does that continue to work as the model that is being fine-tuned becomes better at the task than the frontier model?