I recently switched to using for some large volume inference with small, fine-tuned models and was surprised how nice and painless things were. vLLM seems to have a lot more gotchas and kludged together stuff once you get outside of anything straightforward.