I feel like you would want to run 8 smaller models separately for quantity of raw output. 1 big model is slow and isnt guaranteed to make no mistakes.
Doesn't really mean anything without a specific use case to guide model selection.
The thing is Qwen 3.8 27B can be ran on far far cheaper hardware. If you're spending the big bucks on these rigs you probably made the wrong choice if you aren't using models that require all of that VRAM.
That's not quite how it works. Throwing Deepseek V4 Flash on 4 of these would net you something like >200tk/s for 16 concurrent requests, that's 600 million _output_ tokens a month. Guess what happens when you use 8