It's nice to see time reflected here.
Deepseek is fuckin fast! Seems like the bigger the task, the faster it gets, which is kind of unfortunate because nobody is going to be doing 17 benchmark passes on a $50-100 task. I'm assuming the brief pause before it avalanches out 16kb of text at upwards of 250tk/s (multiples beyond anything resembling a comfortable reading speed) is some sort of workload evaluator directing sessions to individual/multiple cards, occasionally waiting for what it thinks is best to become available.
I couldn't believe it at first, it shit out a damn fine multithreaded physics simulation fabric (integrated into a massive codebase, tests passing) in under an hour. Anything I could find online says they average like 80 but my logs average ~3x that.
DeepSeek is consistently the fastest model, and Kimi the slowest. GPT in the middle. Anthropic is probably on the slow side as well.
>nobody is going to be doing 17 benchmark passes on a $50-100 task
Hopefully we'll see more of this as big companies try to optimize token usage where the cost of benchmarking is dwarfed by the potential savings across the org