That's exactly how they do it.
There are ML models that do the reverse and output image to text, which assist quite a lot.
The better the text represents the unique thing in the photo, the better the model understands what that text means.