logoalt Hacker News

minimaxir • today at 9:53 PM • 0 replies • view on HN

The main one is mapping images to text and visa versa, e.g. semantic search of images via text, where the images are encoded and the text question is encoded with the same model, then finding nearest neighbors.