Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.
It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.
I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware
> Model SOTA moves faster than chips can be designed or produced.
From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252
Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.
I don’t disagree with most of what you’re saying, except for one point: I must have gotten a dumb dog, I’m a little jealous…
They have trillions..
Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.
I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?
At this point, LLM's are "good enough" for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper.
All aboard! We're racing to the bottom now.