It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https://openrouter.ai/qwen/qwen3.8-27b#providers
They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras
We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).
https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.