logoalt Hacker News

overfeedlast Sunday at 12:09 AM2 repliesview on HN

While it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model?

Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.


Replies

ACCount37last Sunday at 8:37 AM

What do you think those tokens are used for?

Distillation attacks aren't about replacing the entire pretraining dataset with questionably sourced synthetics. It's all about post-training.

Train your own base model - but tune it off Claude output to make it perform more in line with Claude. Yoink the products of Anthropic's expensive SFT, RLHF and RLVR work for yourself by training on the outcomes.

The post-training datasets are small, but they are what controls the final model behavior.

show 2 replies
tristanjlast Sunday at 12:40 AM

3.4 million is the number of sessions Anthropic detected. The actual number of Claude sessions trained on is likely >100 million. There are tens of thousands of accounts funneling Claude sessions into Chinese labs https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

They are used for post-training, i.e. calibrating the model to understand and use tools/command line more effectively.

show 1 reply