That "TPU advantage" might be slowing Google down (though likely not as much as their internal bureaucracy).
Porting CUDA-based research, debugging, and overall experimentation speed is likely slower.
The GPU is still king for training.
lmao, you know all Anthropic models are trained on TPU right?
But maybe the TPU advantage is in inference? That's what I assume because the number of compute cycles are going to be all in inference vs training. So they could train on GPUs if they want.