logoalt Hacker News

DFlash 2: Keep Drafting Parallel

68 pointsby mike-the-brainyesterday at 8:28 PM8 commentsview on HN

Comments

ilcyesterday at 11:12 PM

Watch the video carefully. DFlash2's tool call fails on python syntax.

Usually models in this class nail things like that 1 shot, which the other side did.

I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.

show 1 reply
hypferyesterday at 9:22 PM

Amazing tech

> An agent writes in an afternoon what a chatbot writes in a month

But can you just.. not.

Your tech is so good, it speaks for itself. Don't ruin that.

adefayesterday at 9:20 PM

I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.

sarjannyesterday at 9:45 PM

Great news, has made low memory bandwidth model usage so much nicer.

verdvermyesterday at 8:32 PM

vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816

show 1 reply