logoalt Hacker News

Juvination • today at 9:47 PM • 1 reply • view on HN

So what are some use cases people have found for running these sized multimodals on their device? What is it accurate on, and what is the hallucination rate like?


Replies

minimaxir • today at 9:53 PM

The main one is mapping images to text and visa versa, e.g. semantic search of images via text, where the images are encoded and the text question is encoded with the same model, then finding nearest neighbors.