logoalt Hacker News

armcattoday at 1:40 PM1 replyview on HN

How is input token efficiency/verbosity on this model? Has anyone tried? GLM 5.2 was doing lot of turns and thinking piling up input tokens in the context (compared to Claude and GPT models). Then Qwen3.8-27B was 2x of that. Both delivered good output results but those cumulative input token costs were not cheap. Note this is on our specific business workloads. Genuinely interested in other people's experience (if you are able to try it out).


Replies

petutoday at 1:50 PM

Haven't tried, would be surprised if it's any different.

It's new arch demo for future Qwen 4 family, but (as I understand) training recipe/data is same as any other 3.8 model.