logoalt Hacker News

Why isn't the industry freaking out about DeepSeek 4.1 Flash?

269 points • by jonotime • today at 12:14 AM • 233 comments • view on HN

Comments

vishvananda • today at 9:20 PM

The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.

I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.

This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.

➕ show 8 replies
giancarlostoro • today at 7:43 PM

Call me crazy but:

VRAM & Memory Requirements by Precision

• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).

• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).

• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)

VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.

Even so why would anyone not sleep on a model they cannot run?

➕ show 14 replies
profsummergig • today at 10:30 PM

Why isn't the author worried about sending her/his ideas to DeepSeek online (instead of hosting it and using it locally)?

p1necone • today at 8:39 PM

I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).

I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.

However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.

➕ show 1 reply
mlinsey • today at 8:26 PM

I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).

I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).

Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.

➕ show 1 reply
lmf4lol • today at 8:06 PM

Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.

Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.

But as a main driver. I love flash. And it brought our bill down by A LOT :D

➕ show 3 replies
gregwebs • today at 9:42 PM

I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.

DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.

What I use it for is

  * the orchestator of my coding workflows
  * the tester/verifier of code changes
  * the sub agent that explores code or does web searches
  * putting together code base research reports
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.

They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.

Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.

➕ show 1 reply
user43928 • today at 9:23 PM

Because DeepSeek is not "a month or two" behind as claimed in the article.

These open models still did not beat February's Mythos / Fable 5.

DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.

Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.

It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.

➕ show 3 replies
arush15june • today at 9:14 PM

I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.

I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.

Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.

And it never says no for cyber tasks so that's a big win

➕ show 1 reply
hmontazeri • today at 7:55 PM

I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it

➕ show 2 replies
james2doyle • today at 9:43 PM

Been using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness

nerdypepper • today at 10:19 PM

https://tangled.org/astrra.space/ds4-recipe is an incredibly cool writeup on making deepseek v4.1 flash run really fast.

zug_zug • today at 8:56 PM

I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.

That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.

simpaticoder • today at 7:53 PM

The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.

The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.

➕ show 1 reply
wg0 • today at 7:53 PM

While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.

I realized that mistake and guided DeepSeek where it should be.

Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.

➕ show 1 reply
apitman • today at 9:10 PM

> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited

My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).

I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.

This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.

➕ show 1 reply
swiftcoder • today at 7:57 PM

I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them

➕ show 1 reply
RGS1811 • today at 8:24 PM

This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.

I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.

alex-moon • today at 8:28 PM

I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.

aguilaair • today at 8:12 PM

What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.

see https://artificialanalysis.ai/models/releases/comparisons?co...

➕ show 1 reply
potsandpans • today at 10:28 PM

I'm using it quite extensively in my PlayStation decompilation harness

0xbadcafebee • today at 10:28 PM

[delayed]

_jayhack_ • today at 9:54 PM

Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about

ne01 • today at 9:57 PM

Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!

wren6991 • today at 9:50 PM

It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.

elmer2 • today at 8:21 PM

DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.

➕ show 1 reply
jeffrallen • today at 10:25 PM

[delayed]

LeBit • today at 9:46 PM

I have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.

It costs pennies and you got really great output.

The author is spot on.

bitfilped • today at 10:23 PM

Because in two weeks someone will be asking why I'm not freaking out about AlphaDolphins 0.3 Zip and then in a month FrozenMonkey 2.5 Artic.

pants2 • today at 9:31 PM

Probably because Luna is faster, cheaper, and approximately as smart

browningstreet • today at 7:48 PM

What would freaking out look like, or is this just a stupid bloggish title flourish?

Is OpenAI coming in $20B under a sign of "freaking out"?

➕ show 2 replies
smallmancontrov • today at 7:47 PM

They might be. They would delay public admission as long as possible, because public admission would make stocks go down.

try-working • today at 10:17 PM

I have used over 40B tokens and spent over $800 on DeepSeek API over the past 30 days, mostly on V4.1 Flash.

It's good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that's the difference.

f6v • today at 9:05 PM

My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.

jbellis • today at 8:51 PM

I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/

And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking

booi • today at 7:43 PM

Because GLM 5.3 Flash is even cheaper?

➕ show 3 replies
liuliu • today at 8:17 PM

DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.

➕ show 1 reply
wildster • today at 8:48 PM

I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md

➕ show 1 reply
xyzsparetimexyz • today at 8:25 PM

There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.

aussieguy1234 • today at 9:50 PM

What blows me away about this model is it's speed.

It's way faster than Opus or any of the GPT models.

I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).

aszen • today at 8:45 PM

Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out

thefourthchime • today at 7:54 PM

For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.

Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5

Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...

pianopatrick • today at 8:04 PM

I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.

Would be cool if they added it.

➕ show 1 reply
tengbretson • today at 7:43 PM

I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.

gsky • today at 8:10 PM

America bans Chinese models sooner or later just the China banned American big tech

hypfer • today at 8:00 PM

Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?

➕ show 1 reply
pizza234 • today at 8:36 PM

People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.

I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).

Local models are also really slow, unless one spends insane amounts of money.

Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).

➕ show 2 replies
robertlane0 • today at 9:27 PM

Honestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I'm sharing it with Meta a little moot because it's already published permissively on GitHub anyways.

🔗 View 16 more comments