logoalt Hacker News

redox99today at 1:17 AM0 repliesview on HN

I have no problem running two or three sequences of qwen 27B with a 3090. It's basically the recommended way, LLM inference without batching is super inefficient.