logoalt Hacker News

m_ketoday at 2:59 PM0 repliesview on HN

On Policy Self Distillation and Active Learning. Anything that increases sample efficiency by providing a richer more dense feedback signal and is more efficient at exploration / sampling.