logoalt Hacker News

NitpickLawyertoday at 3:50 PM0 repliesview on HN

This is what I copied from the en version of the modelscope page, right when they published it:

> Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

> Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

There was another paragraph about a new attention, but I didn't copy that.