logoalt Hacker News

joefourieryesterday at 11:36 PM4 repliesview on HN

> current frontier models

> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.


Replies

sobelliantoday at 12:40 AM

I tested both myself and a weak bot against Astra xhigh, https://lichess.org/study/27lCQqDa. It's still pretty bad at chess, though it takes longer to devolve into illegal moves.

sigmoid10yesterday at 11:44 PM

The actual current frontier plays somewhere around GM level.

https://chessbench-ai.github.io/#leaderboard

It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.

show 7 replies