I've had 2b models give a plausible Paris vacation itinerary. A tools-capable 12b and especially 30b model from 2026 is certainly capable of producing passable results. I was demonstrating the qwen 3.6 27b model I stood up last week to my wife and it gave her a passable Moroccan Chicken recipe. With tool calling (search) they're quite good.
With OpenAI having released 20 and 120B models a while back, I think they recognized that tiny models were never going to be a defensible income stream.
Any value will come from the largest models, and those largest models are unlikely to ever run on consumer hardware within their window of relevancy.
It’s AI talking about AI so cum grano salis, but my AI is saying I would need at least a half million dollar in hardware to run the newer high quality Chinese models with bemchmark-competitive force.
You can run a Moroccan chicken fragment on the cheap in homage to what you cannot run
I was on a long international flight recently with no internet and Google AI Edge Gallery installed on my phone with a 2.5GB quantized version of Gemma 4 on it. I was able to chat to it for a while and get some information about my destination which all turned out to be true and good advice.