There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
Give away access to the model and go ask people from time to time if the model was of use to the person and if they were able to make the model work with them.
I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it
Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.