logoalt Hacker News

mirekrusintoday at 12:15 AM0 repliesview on HN

Dual 4090, getting 85-113 t/s depending on task (draft seems to speed up quite a lot, disproportionately more for content like svg etc):

  ./llama.cpp/llama-server \
        -hf unsloth/Qwen3.8-27B-GGUF:UD-Q8_K_XL \
        --webui-mcp-proxy \
        --no-mmproj \
        --parallel 1 \
        --kv-unified \
        --flash-attn on \
        --fit off \
        --split-mode tensor \
        -ngl 999 \
        --cache-type-k q8_0 \
        --cache-type-v q8_0 \
        -ub 256 \
        --no-context-shift \
        --host 0.0.0.0 \
        --tools all \
        --jinja \
        --ctx-size 262144 \
        --spec-type draft-mtp \
        --spec-draft-n-max 3 \
        --reasoning on \
        --chat-template-kwargs '{"reasoning_effort":"medium"}' \
        --reasoning-preserve \
        --temp 1.0 \
        --top-p 0.95 \
        --top-k 20 \
        --min-p 0.0 \
        --presence-penalty 0.0 \
        --repeat-penalty 1.0
Use claude/codex/whatever with /goal to optimize params for you.

IMHO draft model support on dense models is great alternative to MoE on GPUs (high bandwidth, less memory) – more intelligence, speed somewhere mid way there which is usually sufficient.