There isn't even deepseek V4. I'd rather trust LLM arena leaderboard, which puts it on p...

aucisson_masque • yesterday at 10:13 PM • 1 reply • view on HN

There isn't even deepseek V4.

I'd rather trust LLM arena leaderboard, which puts it on par with sonnet.

Replies

LM Arena uses human side by side voting, which limits its applicability to complex tasks.

The ARCPrize leaderboard does have Deepseek V3.2, which only scored 4% on ARC-AGI 2 (while the top models score over 80%). It also Kimi and Qwen, but they also didn't perform well.

alt Hacker News

Replies