It’s not just about small models, that’s only one part of evolution
Some groups are baking models into silicone, Deepmind has an example, it gets 18,000 tokens/sec on Llama 3.1, not sure about parameter size
I think this is the future - at least it will be for on-device models. Apple, for instance, will "bake silicon" once a year for their current model, and use that chip in all their devices.
> Some groups are baking models into silicone
While some other groups are baking silicone into models :)