If anything, this report actually made me feel a bit better about AI-led job extinction not being that close. The sheer complexity of the swarm's actions required AI to parse and aggregate, but even with the METR team effectively having an unmetered token budget to do so, the output/summary still required extensive human review.
Even if we ignore the hinted possibility that the agents used to summarize the voluminous data may be acting deceptively (i.e. no snitching), the agents' summaries often 'missed the mark'. Maybe this is another example of 'taste', but it seems less subjective than arguments I've seen for that. It could be that an LLM is no better able to define 'usefulness' or 'relevance' to humans absent being told exactly what that is.
Fair enough, but even so surely putting together the METR report is an example of a highly difficult task that the vast majority of humans are incapable of? It's easy to forget how heavily skewed the crowds on HN and those surrounding engineering and science operations are.