> GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second
> delivering 17k tokens per second per user on Llama 3.1 8B model.
Obviously this is a much smaller model, but I really can't wait for ASICs to take over the LLM space.
Imagine running a model like Sol/Fable (even half the size with 60-70% of it's intelligence) on your own ASIC hardware.
I'd take qwen3.6 (3.8 as of tomorrow) 27B running at 17k per second first on the way to Sol/Fable! And then dsv4-flash-0731!
ASIC makes it sound like it's a single chip, but in reality serving trillion-param models on Cerebras requires a full cluster (as in multiple racks, MW of power).
Some interesting twitter analysis here:
https://x.com/bleysg/status/2073937651150029084