Fine tuning small models is not novel. The novelty is large model generalization without fine tuning, at small models cost/latency.
The OP acknowledged they needed to fine tune their model to the training data of the task vs. zero-shot Jev