logoalt Hacker News

ddp26today at 7:13 PM1 replyview on HN

There must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model.

What's the simplest explanation?


Replies

cogman10today at 7:20 PM

Perhaps post training? I believe I read that Qwen 3.8 is just post trained Qwen 3.6, which is why it was able to be released so quick.

It may be that these flash models are simply post trained larger older models.