> At 1t/s it's still faster than humans for a lot of tasks
Which tasks? I think you're underestimating how token hungry current proposed workflows are.
Anything you do right now? A typical 10 minute prompt "simply" becomes about 7 hours long. (40t/s vs 1t/s).
Doesn't matter which task. Compare it with about 40-50t/s an LLM oneshots with, and it, and whatever task now takes X time, takes X * 40-50 with this.