logoalt Hacker News

moritzwarhiertoday at 3:03 PM0 repliesview on HN

They still scrape code, I'd guess, e.g. from GitHub?

And there's tons of Claude-generated code there.

Also, I'd guess this is not just to prevent any AI-generated code in the training data, but specifically their own.

Percentage of users who put out their code on the web and also have a plan where Anthropic promises not to train on their data is problem also low.

So not excluding own code could be a real issue, since it would be impossible to deduplicate the training and RILHF data from their sessions with the code accessible elsewhere, and written by the very same users.