The cost is 20-40x less for Deepseek Flash v4.1. If you are just comparing to Sonnet or you aren't paying (your case) then your advice makes perfect sense.
I also agree that its a big mistake to have a flash model implement without a strong model reviewing.
I have Opus plan, Deepseek implement the code, and then review with Opus [1]. In this workflow I am saving a lot of money by having Deepseek do the implementation. Note that the review back-and-forth is fully automated [2], so it doesn't take any extra attention from me.
[1] https://github.com/gregwebs/skills-sdlc/tree/main/skills/implement
[2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/code-review-with-followup/SKILL.mdDeepseek Flash v4.1 is only "40X cheaper" if you do not account for the time of the engineer reading the output. If Opus 5.5 high requires 1/2 of the actual engineer time, and the engineer costs $100-$200/hr, then Deepseek v4.1 is actually the more expensive model to use.
I have tried your workflow many times, and simply letting Opus do the implementation costs much less than wasting hundreds of millions of tokens letting deepseek and opus go back and forth and back and forth. And bonus, my project finishes in 5 minutes instead of 20.
> If you are just comparing to Sonnet or you aren't paying (your case) then your advice makes perfect sense.
Also if you aren't hitting capacity.
Off-work, I use LLMs regularly for both design/coding and non-technical work, but the volume is not enough to trip the weekly limits, and rarely enough to trip the daily limits. So I just go with whatever's current best SOTA available on my Claude & ChatGPT subscriptions and don't worry about limits. If I hit one, I do some household stuff or relax for a few hours (or just turn in for the day), and then the limit is refreshed.