logoalt Hacker News

umviyesterday at 11:30 PM4 repliesview on HN

I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.

There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.


Replies

jessetempyesterday at 11:38 PM

All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias

hadlockyesterday at 11:36 PM

I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector

show 1 reply
satvikpendemtoday at 1:06 AM

This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.

asawfofortoday at 12:44 AM

Isn’t this what the paper-bound encyclopedia companies do, albeit shallowly