logoalt Hacker News

anerli • today at 6:25 PM • 1 reply • view on HN

Yeah these are all things that we directly tackle!

Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).

Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.

Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.


Replies

skohan • today at 6:47 PM

Do you have anything published on the quality benchmarking using your caching strategy?

➕ show 1 reply