logoalt Hacker News

birdatlaw • today at 5:03 PM • 1 reply • view on HN

From what I've read, not only are some labs doing it (other commenters already mentioned).

But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.

I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.

I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.

I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.


Replies

fsiefken • today at 8:45 PM

I take 200W for 17,000 tokens/sec any day, even just for Llama 3.1 8B. For that size I want to see this one on silicon. https://huggingface.co/xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Te...