Very curious to see how this compares to Deepseek v4 Flash. I have to assume they wouldn't be releasing this if it was worse.
Their "next" variants are usually undercooked, but useful for the community to verify support for inference stacks. This will likely be the same.
Why not? It's not really competing in the same size class.
Besides, as they explicitly wrote here, the main goal for this release is not performance, rather to serve as a reference for inference runtimes about what to implement. So that later Qwen 4 can be released with zero day support.