I find them useful bellweathers of genuinely out of domain performance and capability, regardless of their theoretical importance. What I see is that performance trends are remarkably stable both upstream (miraculous scaling laws of pretraining on validation loss) and downstream performance (epoch capability index). We get the equivalent of a GPT4->GPT5 performance leap every ~16-18 months, and we are not hitting ceilings nor do we see any deceleration.
Today we can solve nontrivial open problems. What will we be able to do next year or the year after? 18months ago no one was using a coding agent seriously. Now for a large segment of the population you cannot do your job without them.