logoalt Hacker News

calebkaiser • today at 2:30 AM • 0 replies • view on HN

This is a productive line of research with lots to explore, if you're interested. There are different flavors based on what you're describing that you might like. Maybe you already know all of this, but just in case any other reader is curious :)

A classic paper on "expanding" a neural network (the second part of your distil-then-grow loop) is Net2Net. Their goal is basically to get a smaller model's knowledge transferred into a larger one to bootstrap the larger model training: https://arxiv.org/abs/1511.05641 . Bert2Bert is basically the same idea but language models: https://arxiv.org/abs/2110.07143

There's also stuff on "growing" a pretrained LLM that is somewhat similar to what you're suggesting: https://arxiv.org/abs/2303.00980

And then there's a bunch of stuff on "refreshing" weights in the network as well, sort of like pruning but without reducing capacity.

One of the things to explore are assumptions around how information really compresses down in distillation, and what is really happening when we deem a model saturated. There's often an intuition that when you distill the model down, you're finding a sort of "truer", more essential representation of the functions you're modeling. And so if you distill it down and then expand the weights, you've added additional capacity for it to learn more, since the distilled core now contains the important information from the original larger model. But very often that's not actually how it plays out, and the problem is thornier than it seems on its face. Plateaus in training don't necessarily mean the model has run out of representational capacity. And distillation is not guaranteed to preserve the most valuable information. It's entirely possible that the representations the original model has learned for performing on the tasks it is evaluated against are more generalizable than the compressed representations you've distilled.

Initialization is also tricky when you're expanding the weights like this. Non-linearities make it such that splitting the weights directly like you're describing doesn't preserve the function. Not that this means the model can't recover from further training, but it's just not where you want to start training from if you can avoid it. There's also a strong likelihood that the added parameters are redundant relative to the distilled core, which will make learning more challenging. A lot of work in this space as a result focuses on stuff like symmetry breaking. A cool paper you might find interesting that isn't about LLMs, but has a similar "Progressively give the model more degrees of freedom" flavor is ARC-AGI Without Pretraining: https://iliao2345.github.io/blog_posts/arc_agi_without_pretr...

But there's lots of active research around this general space. You should test out your ideas and share them! I can't recall now exactly but I feel like I've seen MoE stuff that touches on some similar ideas.