logoalt Hacker News

rao-vtoday at 12:33 PM1 replyview on HN

Native dflash support on day 1 helps a lot! High quality speculative decoding speeds up a lot of agentic work.


Replies

cmrdporcupinetoday at 2:08 PM

You're right. I'm getting ~33tok/sec w/ dflash on it, even bursts up to 60tok/sec, using my personal home-built-for-Spark inference engine (not vLLM or llama.cpp based)

That's pretty respectable.

Still working on optimizing and cleaning up before I push it.