logoalt Hacker News

petuyesterday at 7:30 PM0 repliesview on HN

I have no idea, but I've assumed that batching can't work on Cerebras.

Batching works because of severe memory bottleneck, but Cerebras whole thing is serving models out of "L1 cache" (?).