Qwen has set an excellent track record for architecting and releasing open-weight models that consumer-grade devices can run. What is needed the most right now is something similar to Bonsai 27B, with a modest memory footprint, but faster and more capable. On-device models can make up for intelligence by being faster, thinking longer, or doing more quick iteration rounds.
I’d like a “Bonsai 2.8T.” That is, something that is near the Fable/Sol/K3 class, but capable of running locally on consumer hardware.
I can't really blame them that the biggest labs focused on trainig and realeasing huge models.
The niche for small models should be filled with medium sized labs doing distillations of the huge ones into consumer grade hardware runnable models and LORAs for the huge ones.
> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable
Yeah, that'd be neat, but that's not what this announcement is about at all:
> With a massive 2.4T parameters