logoalt Hacker News

harimau777yesterday at 11:44 PM1 replyview on HN

It seems like an LLM potentially could learn that way if each practice game it participated in was added to its training data.


Replies

willmarchtoday at 12:43 AM

Yes, this is essentially how AlphaGo and AlphaZero algorithms work to train superhuman Go/chess/shogi agents. It’s an elegant algorithm that is analogous to how humans learn games.

show 1 reply