logoalt Hacker News

DeepSeek v4.1 Flash Is Now Our Best Hacking Model

139 pointsby talhof8today at 12:19 PM56 commentsview on HN

Comments

gertlabstoday at 5:29 PM

We ran v4.1 Flash through our evaluations and found it to be smarter and faster than V4 Flash, with a commensurate price bump. Some notes:

- Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice).

- Closer than expected to the top open weights models (GLM 5.3 and Kimi K3) in agentic coding, at lower cost.

- Chinese models have always been strong iterators in an agentic harness. This model is no different, reaching an average percentile ~20% higher when given a harness vs a one-shot solution. That one-shot fluid intelligence is what makes a model feel smart, though, and typically results in fewer attempts/tokens to solve a problem, and American frontier models are still far ahead in that department.

The new architecture is interesting. It puts pricing between their old Flash and Pro lineups, suggesting they might be abandoning their super-cheap flash models (which weren't that fast due to heavy reasoning) and their pro models (which sort of flopped and weren't consistently better than their flash models, despite the size/cost) and shipping a strong intermediate that competes with the Gemini Flash series.

Data at https://gertlabs.com/rankings

habosatoday at 4:45 PM

DeepSeek models have such good benchmark performance, amazing pricing, and the team over there seems to be widely considered impressive.

I just haven't found them to be very good? I've had a ton more success with the GLM models (since 5.2 anyway). Maybe I'm just holding it wrong, DS models seem to get stuck in loops or tell me nonsense. GLM feels like budget Claude.

show 6 replies
TuxSHtoday at 3:15 PM

I find this - or perhaps the title - a bit surprising.

I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.

Perhaps DS works better where targets have low-hanging fruits than can be found fast?

show 10 replies
jrflotoday at 3:31 PM

Seems pretty bold to claim deepseek is the "best hacking model" while providing zero comparisons to other models...

show 2 replies
pelzatessatoday at 6:10 PM

How do I make deepseek "hack" my source code? do I just start my coding agent in my directory and command it to "find vulnerabilities", or is there some more sophisticated software to do that?

hactuallytoday at 6:54 PM

we recently got this running in 192gb of vRAM and using it with the Klaudia harness has been incredible for driving out work that we'd need to use Opus and Fable for previously

Art9681today at 4:44 PM

As opposed to what? Is enclave.ai signed up for GPT Cyber or Glasswing?

fratoday at 5:04 PM

Interesting result, but the writing is very poor.

wg0today at 3:48 PM

DeepSeek is underrated. Basically all Chinese models are good enough for day to day coding at this point.

The 2 trillion dollar ROI on anthropic alone?

Good luck with that.

show 1 reply
fwiptoday at 3:00 PM

> The accepted runs cost only $4.65.

Estimating the cost of unaccepted runs (those that were not successful?) is left as an exercise for the reader.

nickysielickitoday at 3:27 PM

When the history books are written and all is said and done, the hubris of this moment where all the American labs decided to punk their investors and join hand in hand in agreeing to let the Chinese win forever is going to be the main story.

show 2 replies
strictneintoday at 4:48 PM

[dead]