Given the huge amount of money being spent on AI chips in the US, what prevents US AI labs from doing the same level of software optimization? It could be a solve for some of the capacity constraints.
They do, when Luna got 5x cheaper it was directly attributed to some unknown % inference optimization.
US labs are quite cut throat about dealing with stuff costing them money (inference). This sort of engineering excellence doesn't always feel that way because they are simultaneously quite lax about stuff costing other people money.
> what prevents US AI labs from doing the same level of software optimization?
Because they don’t have to. Most of the time money would buy you newest and/or more hardwares so there’s low/minimal interest to optimize the code or approach.
They have already been doing it for months https://openai.com/index/openai-broadcom-jalapeno-inference-... . OpenAI on their custom chip brought up lightspeed deepseek as experiment by using AI in the exact same way as this zAI blogpost. And the kernel optimization contests/etc have all been havily done through AI based optimization loops for half a year+.