logoalt Hacker News

CMaytoday at 5:07 PM0 repliesview on HN

> if you run one model, run glm-5.3

That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.

Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.

Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.