What protection do LLM search engines have against training off content generated by other LLMs?
Will we get to a point where AI-generated sites make up a majority of the internet, and LLMs are training upon their own regurgitations, with exponential amplification of all their lies and flaws?
Or will the pre-2022 corpus human knowledge be considered the low-background steel standard, and anything after that less and less reliable unless certified that it has been created by a human mind and untainted by hallucinations?
They’ll train on prompts and anything else you send in. Many LLM responses are sorta finger printable: I assume this is intentional
I've mostly stopped using the Internet to learn new things and have gone back to books from the library. The majority of technical books at the library were published pre-2020s and hopefully, publishing slop physically won't be profitable enough to flood that market, too. Now that the Internet has largely been destroyed by slop manufacturers, whether or not the words are(/were) worth putting on paper becomes a useful discriminator.
> What protection do LLM search engines have against training off content generated by other LLMs?
You're talking about a scenario that won't blow itself up in the next few quarters, so it's of no interest to them.
I have a feeling we're already there.