> careful layering of well-understood optimizations—RoPE, SwiGLU, GQA, MoE
They basically cloned Qwen3 on that, before adding the few tweaks you mention afterwards.
> They basically cloned Qwen3 on that
Oh, come on! GPT4 was rumoured to be an MoE well before Qwen even started releasing models. oAI didn't have to "clone" anything.
You seem to be conflating when you first heard about those techniques and when they first appeared. None of those techniques were first seen in Qwen, nor this specific combination of techniques.