Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.
Lead times are so long that there is a lot of risk the chips would be obsolete by the time they come out.
Also, it's hard to get fab capacity for any project. Let alone something so experimental.
Yes, and startups do exactly that. Check out Etched https://www.etched.com/ where they made a Transformer specific GPU (basically a form of ASIC) where they bet that transformers would be the dominant GPU architecture for running AI / LLM workload.
Most responses here are along the lines of "model capabilites move too fast to build hardware for".
I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.
From what I've read, not only are some labs doing it (other commenters already mentioned).
But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.
I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.
I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.
I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.
Etched is a startup doing exactly this.
The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model.
So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.
Cuz once they're in chips they can be put into robots, and once they're in robots it won't be so easy to reach the off switch, and once we can't easily reach the off switch, we're doomed.
This question would be better answered if people were careful about distinguishing between "models" and "transformer architecture".
If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs.
Would you pay to crystalize one of today's models in silicon so you can use it in 2028, or would you wait for another 6 months to see how models improve before pulling the trigger on that kind of commitment?
We are still in the middle of the AI race. Commodity hardware is easy to use, can do everything and is fast enough.
Your optimized hardware chip might be obsolete before its back from the fab.
SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.
Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.
Models fully deprecate in a few months. Why would you burn an algorithm that fully depreciates in value faster than a bag of potato chips. The 'inefficient' general purpose hardware is constantly renewed with every released model. Even 6 year old Ampere GPUs are still usable.
Because the iteration speed on models is so fast that by the time they have an ASIC ready for one model version, they are already significantly ahead in capability. Think of how big the jump between Opus 4.8 and 5.5 has been. They were released four months apart.
taalas did: https://taalas.com/
I'd presume because it take too long to go from design to tapeout to production. Their whole business is predicated on having better models.
Also can't keep them closed source if you do that.
Yeah, and what about FPGA? Which was the same interim state when Bitcoin went GPU -> FPGA -> custom chip fab?
Not sure that's actually practical at the scale of SOTA models.
They do. It takes time to deploy those chips though. Check out OpenAI and Broadcom deal.
Even dumber question: What is new or novel about this openTPU?
In addition to some of the other replies you got, here is one more:
Much of a model are weights, and high-density ROMs are very very very hard.
[flagged]
Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.
It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.
I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware