logoalt Hacker News

gentiletoday at 6:54 PM0 repliesview on HN

It's still typically for code, like python data stuff, bash scripts, general web search, small javascript stuff for my website. By saying not SWE, I mean I don't really see much benefit from "agentic" stuff, although I've tried. 50 t/s means 40~60 seconds for a typical thinking response. I used to run gemma 4 e4b-it-qat fully in GPU (~150 t/s), but the quality improvement moving to a much larger MoE model was 100% worth the switch. Especially because I had a lot of idle ram (from the before times :( ). There is some tradeoff in terms of context length, but I just use it as a chat. Also it's fairly trivial to setup as long as the GPU is supported.