logoalt Hacker News

jadboxtoday at 5:25 PM2 repliesview on HN

I need someone to run actual benchmarks between the two.


Replies

NitpickLawyertoday at 6:26 PM

Only relevant benchmarks are those you make yourself, targeted specifically for your workflows. Anything else is just number go up on a pretty graph, and every model out there is probably benchmaxxed to hell on the public ones anyway. Keep yours private.

swatcodertoday at 5:36 PM

Benchmarks are the BMI of model evaluation.

They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.