Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
[flagged]
Maybe taking this up one step to Hessian matrix estimation would improve the asymptotic speed.