logoalt Hacker News

cmrdporcupineyesterday at 8:38 PM1 replyview on HN

Probably. I've spent zero time with optimization at this point. Code is all new this morning.

Curious if llama.cpp is doing full BF16 for the n-gram embeddings table, or if that's quantized, too. I was going to try nvfp4 for that but wasn't sure about the quality risk.

Spark usually wins on prefill, not decode. I think memory bandwidth is about same between the two. What do you get for prefill/prompt?


Replies

hedgehogyesterday at 8:58 PM

Ok, at 50k context its about 126 prefill, 13 generation.

show 1 reply