logoalt Hacker News

willmarchtoday at 12:43 AM1 replyview on HN

Yes, this is essentially how AlphaGo and AlphaZero algorithms work to train superhuman Go/chess/shogi agents. It’s an elegant algorithm that is analogous to how humans learn games.


Replies

zug_zugtoday at 2:18 AM

Well except AlphaZero played 44 million chess games in that time (and actually played with a 44 core computer). So I'd like to point out that the human is still just a few orders of magnitude more efficient.

show 1 reply