How can it be that a 438B model is worse than Qwen3.8-27B? Are these benchmarks totally gamed?
Training data plays a huge role. As an example, Qwen 3 was generally considered a significant improvement over Qwen 2.5, but the architecture only had minor tweaks. The major change was the quantity and quality of training data they used for Qwen 3.
qwen is a magical model. it has the mandate of heaven. was so already at 3.6-27B
Training data plays a huge role. As an example, Qwen 3 was generally considered a significant improvement over Qwen 2.5, but the architecture only had minor tweaks. The major change was the quantity and quality of training data they used for Qwen 3.