How would you filter out garbage from your training data, for example? If you are trying to use someone elses corpus, would it be "from the scratch" then?