If training and inference is hardware constrained, and you can train and server better and bigger models on the same hardware with memory optimizations, that's exactly what I would expect companies to do.