logoalt Hacker News

Dust: Pretraining Transformers Without Backpropagation

53 points • by E-Reverance • today at 9:15 PM • 4 comments • view on HN

Comments

whatshisface • today at 10:51 PM

Maybe taking this up one step to Hessian matrix estimation would improve the asymptotic speed.

polyomino • today at 10:27 PM

Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory

api • today at 10:03 PM

It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?

➕ show 1 reply
derin-picment • today at 10:03 PM

[flagged]