logoalt Hacker News

XenophileJKOyesterday at 11:50 PM1 replyview on HN

It is even worse.. This is a classical reinforcement problem where data generation is easy because the rule set is pre-defined. So you really don't even need any data to start with (but would help).


Replies

manquertoday at 12:12 AM

There are more possible game combinations than atoms in the universe, even those generation of valid game states are as you say pre-defined. that is why models cannot go this route and therefore are poor at chess

show 1 reply