> Any benchmarks using it?
A challenge as I understand it is reproducibility.
Normal LLM runtimes aren't typically fully reproducible even with same random seeds for distribution sampling, due to floating-point numbers, batching and such.
Though averaging over many runs could alleviate that I suppose.
While it would measure some aspects of intelligence, I'd argue it fails to capture other, more creative aspects.