logoalt Hacker News

cmrdporcupinetoday at 5:44 PM0 repliesview on HN

I have nvfp4 quant fitting fine in 128GB on DGX Spark, but with paging (from nVME) of the n-gram table. Resident ~80GiB for weights & context.

On branch of https://github.com/rdaum/eider (for DGX Spark). ~12 tok/sec decode without speculative decoding (will come later)

Still actively working on this. Prefill currently sucks. Will merge to main by end of day.

EDIT: This has now landed on main. Still haven't done MTP speculative decoding boost, but:

80tok/sec prefill, 12 tok/sec decode. ~90GiB or so resident. n-grams paged from disk.