logoalt Hacker News

applicativetoday at 2:40 PM1 replyview on HN

So they do it as a free service to the other LLMs?


Replies

moritzwarhiertoday at 3:03 PM

They still scrape code, I'd guess, e.g. from GitHub?

And there's tons of Claude-generated code there.

Also, I'd guess this is not just to prevent any AI-generated code in the training data, but specifically their own.

Percentage of users who put out their code on the web and also have a plan where Anthropic promises not to train on their data is problem also low.

So not excluding own code could be a real issue, since it would be impossible to deduplicate the training and RILHF data from their sessions with the code accessible elsewhere, and written by the very same users.