Use as training data has been tested in court multiple times and cleared.
I know some people want it to be a legal violation, but it’s not.
Also these are audio models, if you hadn’t noticed. The training sets for this type of work is very different than the corpus of scraped GitHub repos.