> How do you measure that?
Vibes.
Reviews take way more time than new features, and often utilize 3 subagents at a time - it's not uncommon to have reviews take 3 hours when the feature only took 15 minutes. The numbers here are just estimates, though - I don't keep track. "10x to 20x tokens on review" is probably underestimating review cost, if anything.
I don't review all code myself, but I do spend down-time reading code. When I find obviously bad code/patterns/whatever, I don't just fix it directly - I work to create automated tests (like static analysis, almost) or adjust or create prompts and tools for agents (either for review stage or initial implementation) to solve that class of issue, instead of just that instance of it.
I never review the first iteration. I review it after the agents have already poured a ton of time into it, if ever.
I do sometimes watch agents work, and redirect them if I see them doing something dumb. But, usually, I have 5+ agents running, not counting sub-agents, and I'm just messing with the apps myself to find the next feature to create or polish.