Reading their paper, it wasn't trained from scratch, it's a fine tune of a Qwen3-32B model...

kevmo314 • yesterday at 8:24 PM • 0 replies • view on HN

Reading their paper, it wasn't trained from scratch, it's a fine tune of a Qwen3-32B model. I think this approach is correct, but it does mean that only a subset of the training data is really open.

alt Hacker News