logoalt Hacker News

anerli • today at 6:18 PM • 0 replies • view on HN

The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.

For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.

We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.

The source for the benchmark is available here also: https://github.com/magnitudedev/magnitude/tree/main/inferenc...