logoalt Hacker News

manofmanysmilestoday at 7:03 PM2 repliesview on HN

Imagine this, and sucesor models on Cerebras or other silicon...


Replies

WASDxtoday at 8:01 PM

That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology.

If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.

Moduketoday at 8:25 PM

Very exciting indeed. It is in the works. Their current dense offering, Gemma 4 31B, sits at ~1800t/s

https://news.ycombinator.com/item?id=49308715