logoalt Hacker News

criley2 • today at 1:33 PM • 2 replies • view on HN

Deepseek Flash v4.1 is only "40X cheaper" if you do not account for the time of the engineer reading the output. If Opus 5.5 high requires 1/2 of the actual engineer time, and the engineer costs $100-$200/hr, then Deepseek v4.1 is actually the more expensive model to use.

I have tried your workflow many times, and simply letting Opus do the implementation costs much less than wasting hundreds of millions of tokens letting deepseek and opus go back and forth and back and forth. And bonus, my project finishes in 5 minutes instead of 20.


Replies

gregwebs • today at 2:10 PM

I tried the experiment of reviewing vs. not reviewing with frontier models. I consistently found that reviewing by a model with independent context finds important issues when changes are non-trivial- certainly the definition of non-trivial is getting raise as the models get better.

I do have an /implement-simple workflow to skip the planning phase, but even that doesn't skip the review.

Are you doing your own intensive reviews of the model code? Can you share the prompts you are using as I have?

My bar for what models produce without human intervention is much lower defect than what a human would produce. The human interaction is mostly to guide the design and then the review burden is very low. I suspect your bar for what agents produce is lower- you are taking more of the review burden. I also suspect that you are measuring time more than actual cost since your employer is paying and that you are comparing to Sonnet rather than DeepSeek (DeepSeek 4.1 again is 20-40x cheaper than Sonnet). You mention hundreds of millions of tokens (my reviews don't use that much), but even that costs ~$1 on the DeepSeek side.

I think you are taking exactly the right approach at your employer given the cost is free and you only have access to Anthropic models.

➕ show 2 replies
Roark66 • today at 2:07 PM

Bingo, DeepSeek (v4.1) is horribly overhyped. In all my personal benchmarks it sits below Glm5.3 Flash. Waaaay below Qwen3.8-Flash-Next a model less than half it's size.

No, the only open weight model that really makes sense for me is Qwen3.8-Flash-Next, but it is mainly because I can run it locally with reasonable speed (prefill between 650-1400t/s generation between 22-50t/s depending on number of slots/users I configure).

This is the first model that truly competes with Opus 4.8. I'd say it may be better than Opus 4.6 on programming.

But it is very verbose when it comes to reasoning tokens. The more difficult the task the more verbose it is. Certain very hard tasks that take opus 4.8 400k tokens take Qwen3.8-Flash-Next 2M tokens... But it finishes them.

And what you loose on the generation speed you get back on input caching you can keep on for weeks.

It really depends on the workload.