dont we all deem the ability to improve large models as the defacto capability to produce small ones?
I don't think so, they have different constraints and require different optimizations, being able to produce one of them doesn't mean you'll automagically be good at the other.
I don't think so, they have different constraints and require different optimizations, being able to produce one of them doesn't mean you'll automagically be good at the other.