logoalt Hacker News

epolanskiyesterday at 10:20 PM4 repliesview on HN

I don't get the point of these benchmarks, what are they supposed to represent practically?


Replies

acdhatoday at 1:55 AM

The vendors like to toss around terms like “thinking” or “reasoning” to encourage prospective buyers to anthropomorphize models. These challenges are a visceral reminder that none of those marketing claims accurately describe was LLMs do: they’ll happily return things even the worst human illustrator would never hand in and make errors showing that there’s no model of the world behind anything they do.

That doesn’t mean there are no ways to use them productively but rather that you should keep in mind that the same model will happily give you code or a decision with the same level of error unless you have carefully setup a QA regimen to prevent that.

Mindless2112today at 2:29 AM

Unless the LLM has an SVG renderer in its toolbox, it's like asking a human to draw while blindfolded. It's amazing if they can do it well, but it's practically meaningless.

viccisyesterday at 10:24 PM

The ability of an LLM to produce something not in its training data set.

show 1 reply
runarbergyesterday at 10:25 PM

For me this looks ideological (or even political), not practical. The theory is that LLMs are approaching general intelligence (whatever that means) and that the more generic of a task they can perform—no matter how badly—the closer we are to AGI.

Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.

show 1 reply