logoalt Hacker News

yolandac • today at 5:59 PM • 1 reply • view on HN

does it allow us to run larger models that weren't possible before?


Replies

anerli • today at 6:49 PM

Right now, since we use less memory for KV, you have more room for model weights when you're running longer sessions.

However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.