Distributed training is much harder than distributed inference but not impossible. See the recent development of DiLoCo at Nous Research and Prime Intellect.