logoalt Hacker News

porneltoday at 3:12 PM0 repliesview on HN

IANAMLE, but there is "grokking" that makes models learn to actually generalize, even after you give them enough parameters that would let them memorize the dataset:

https://en.wikipedia.org/wiki/Grokking_(machine_learning)

High-dimensional gradient descent behaves very differently than the simplified 3d visualisations we use to demonstrate it, and has lots of ways out of local minima:

https://www.youtube.com/watch?v=NrO20Jb-hy0

so it seems like there is a benefit to giving models more space to learn in rather than forcing them to compress the knowledge from the start.