I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.
There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.
Isn’t this what the paper-bound encyclopedia companies do, albeit shallowly
All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias