"SWE-2 is post-trained from Kimi K3"
I presume post training is significantly easier than the distillation/training the top Chinese labs are doing.
I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.
Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.
Maybe still worth it if their "64% cheaper" figure holds.