The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.
Ah, the old "good, fast, or cheap; pick two" proves true once again.
Some of us want fast food
Not if your use case needs speed. For one of my products I can't use an LLM that has a p99 of >700ms for TTFT.
It means something, because it an iterative workflow. If you're willing to burn tokens, it's possible for weaker models to implement tasks by incrementally improving drafts.