I think one of the end-game components for AI is to be able to read all the literature and generate a reliable list of which ones are not correct and a convincing reason why.
> generate a reliable list of which ones are not correct and a convincing reason why
That is not realistic, but I suppose where things are heading is that you have some indicator of the strength of evidence -- see Fig. 13 in the following insightful take:
https://news.ycombinator.com/item?id=49407226
Though, "strength" should probably be "reliability" and "validity", and I suppose those indicators are more for picking signals from the noise; i.e., what is even worth clicking and reading. That would be increasingly valuable already today due to the volume (and, yes, slop and other related stuff).
>... and a convincing reason why.
I believe this is one of the LLMs' more destructive elements. The more effective ones are quite good at providing output that gets the meat to repeat the output to other meat. Whether the output actually benefits the meat is disconnected from whether or not the meat accepts and repeats it.