It's pretty close to how we measure IQ. The standard test is basically a series of puzzles.
I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.
Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.
Pure software benchmarks might be getting saturated, but physical ones aren't.
Let LLM control a physical robot to perform tasks that average human can do.
Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.
I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there.
It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.