logoalt Hacker News

tristanMatthiastoday at 8:31 PM2 repliesview on HN

> GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second

https://taalas.com/products/

> delivering 17k tokens per second per user on Llama 3.1 8B model.

Obviously this is a much smaller model, but I really can't wait for ASICs to take over the LLM space.

Imagine running a model like Sol/Fable (even half the size with 60-70% of it's intelligence) on your own ASIC hardware.


Replies

mNovaktoday at 8:59 PM

ASIC makes it sound like it's a single chip, but in reality serving trillion-param models on Cerebras requires a full cluster (as in multiple racks, MW of power).

Some interesting twitter analysis here:

https://x.com/bleysg/status/2073937651150029084

auspivtoday at 8:35 PM

I'd take qwen3.6 (3.8 as of tomorrow) 27B running at 17k per second first on the way to Sol/Fable! And then dsv4-flash-0731!

show 2 replies