On my M3 mac, it works okay inside a docker container with just CPU.
{ "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 }
This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292
There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.