fine-tuning may be a more scalable approach to LLM personalization than sending all the same context to two LLMs
I'm working towards both in my homelab to see which works better with little qwen