Just out of curiosity, why run "local-local" when you could just set up a Mini or Studio at home and query it over http? [edit] whole conversation about this in another thread https://news.ycombinator.com/item?id=49433413
I’m personally considering retiring my MBP for a Studio + 15" Air whenever this MBP ages out.
This is the way.
I’m doing that. Mac Mini M4 Pro with 48G RAM as a headless llama.cpp server.
I much prefer using " thin clients " as the interface to the big VMs running in my homelab