logoalt Hacker News

ahmedhossamdevtoday at 3:48 PM0 repliesview on HN

The replay simulator from history for off-policy eval is clever - avoids expensive rollouts. Curious how they prevent the policy from overfitting to already-discovered branches and going stale as the search space expands?