logoalt Hacker News

alyxyatoday at 1:21 AM0 repliesview on HN

Fundamentally I don't believe second-order methods get better data efficiency by itself, but changes to the optimizer can because the convergence behavior changes. ML theory lags behind the results in practice.