logoalt Hacker News

fragmedetoday at 8:24 AM1 replyview on HN

No it isn't. Best and worst and ill-defined anyway but the chess ELO score of various LLMs has fluctuated up and down, it's not been montonically increasing. What is the best answer to "how do I make cocaine"? The models are getting larger, with more compute and RAM backing them, but that doesn't automatically make them better if you don't define how you're measuring better-ness.


Replies

famouswafflestoday at 2:41 PM

None of the frontier labs care about Chess as it's already a solved problem. If they did, the models would be much better. It's really not that hard. Google has a paper on grandmaster level chess without search from transformers.

Better obviously means better, like how they became better than they were 6 months and a year ago.

show 1 reply