GPU failures are frequent enough that at a certain scale, you constantly have workers dropping out. Designing systems that can still keep training is very interesting!