Came for: "A computer once beat me at chess, but it was no match for me at kick boxing."
TFA was actually about leaps of intuition, sadly.
One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
This is an interesting experiment but I wonder if it would be possible to prevent some sort of retrospective bias. For example, I’d expect the experiments that lead to relativity to be over-represented in our catalogue of scientific literature prior to 1900, just because in retrospect they were important, so the records about them were preserved.It would have to be a very intentionally constructed corpus, I think.
I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now
If would be interesting to see 5 billion LLM's working together, each with random mutations (temperature ig). Would we essentially be looking at a society through a petri dish? Ofc 5 billion is quite a lot of compute.
The curious case here is how much of a description do we give it of itself? That would almost certainly dominate success rates.
My feeling is that a prompt would have to provide a vague description of a program that meaningfully passes something like a Turing test, an API to conform to, an expectation of novel construction (no 'ifs all the way down'), and then a requirement to search broadly and pursue promising ideas and not get hung up on the philosophy. Anything more precise feels like it would corrupt the test, but as it is that description feels doomed to loop before even trying the interesting parts.
Article was plenty interesting to me.
> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
Could be, but preventing leakage from more modern stuff can be challenging.
This was attempted with Victorian public domain content: https://www.estragon.news/mr-chatterbox-or-the-modern-promet...
I can't find the citation right now, but I think people found it was leaking anachronisms? So this probably wasn't as well filtered as the creator had hoped?