logoalt Hacker News

gitpusher42yesterday at 9:59 PM0 repliesview on HN

Yeah, I checked it. One expert is about a 3.36mb block. If a cache miss happens I read whole block with one pread.

And there is some reuse. ~41% selected again for the next token, ~57% within two. Each layer has its own experts, so no reuse between these layers.