logoalt Hacker News

blfrtoday at 7:12 AM2 repliesview on HN

I don't care about benchmarks. Benchmarks show that Opus 5 is a stronger model than Fable 5 which is obviously not the case.

But I do care about capability and so far only Anthropic and, very recently with Astra, OpenAI can deliver on coding quality. And capability matters immensely. There is a world of difference between being able to do something and not being able.


Replies

zorkedtoday at 7:24 AM

People have been using LLMs for two year. It's not just this week's LLM release that is capable something.

show 2 replies
mdp2021today at 8:02 AM

> don't care about benchmarks

You must care about good benchmarks (identify those that have relevance).

show 1 reply