From the METR report:
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
No human could have read the reasoning traces by themselves:
> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.
No human could have read the reasoning traces by themselves:
> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.