> Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus.
How does this suggest anyting of the sorts?
> moved by +20.58 points toward GPT-5.5
Score go up. Probability go up. Conclusion.
Appendix B of https://stolen-thoughts.com/paper.pdf discusses this