logoalt Hacker News

lucrbviyesterday at 8:55 PM0 repliesview on HN

They are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.