I sense the approach overcomplicates things. I wonder how long this would have taken in a single session. Maybe this is actually a textbook example of something where subagents make sense, but when I first started LLMs I was often overcomplicating the workflow with all kinds of orchestration. Now I just use one agent, it better allows controlling the output even if the agent works slightly longer. Most time is spent by me writing prompts and reviewing work anyway (for me at least).
The best models have O(1k tokens per second). A billion tokens, sequentially, takes 1-2 weeks. 500B takes 500x that ... not fast. The problem might not have required 500B tokens, but if it needed anything within a couple orders of magnitude then something like the given approach (or anything else yielding equivalent results in exchange for parallelism) was mandatory.