> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
Could be, but preventing leakage from more modern stuff can be challenging.
This was attempted with Victorian public domain content: https://www.estragon.news/mr-chatterbox-or-the-modern-promet...
I can't find the citation right now, but I think people found it was leaking anachronisms? So this probably wasn't as well filtered as the creator had hoped?
Progress followed improvements in hardware, would you have to give access to modern hardware in the experiment for it to use? How much could it infer from it?