An adjacent question: is there an input dataset you can use for training that be computed in closed form so that when you train on your target dataset, learning is effecient.
Methods like formula driven supervised learning exist to arrive a good pretrained weight state, but could this procedure be generalized for specific datasets or flavors of input data.
reminds me of perturbation theory -- start off with a nearby problem you know the answer to, then update it to get the answer to the problem at hand