Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.
you can probably generate quite a few example pairs in a single shot, you also likely don't need the best models for this either
Good observation! It would have to be offset with O(140k) queries to the model, which is, well, unlikely.