logoalt Hacker News

Aurornis • yesterday at 8:08 PM • 2 replies • view on HN

A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.

For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.

The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.

Think of it as distillation, but focused on a specific task.


Replies

nearbuy • yesterday at 8:26 PM

Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.

➕ show 2 replies
jamienk • yesterday at 8:46 PM

This is so cool - I'm aware of this in a vague way. Can you write a little tutorial or give some good links. I want this to be the next new things I do :)

➕ show 1 reply