logoalt Hacker News

pietztoday at 4:03 PM1 replyview on HN

With tiny models surpassing huge, 6 months old models on benchmarks, does anybody have some smart words to share on how these still "feel" different?

Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.


Replies

KptMarchewatoday at 4:27 PM

I agree. They are definitely good - no issues with instruction following for example - but they miss the "intelligence" larger models have.

For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.

show 1 reply