I built this same solution for myself last year, used Hugging Face's "SmolVLM". It wo...

spacecadet • yesterday at 8:29 PM • 0 replies • view on HN

I built this same solution for myself last year, used Hugging Face's "SmolVLM". It works surprisingly well. I use the model to generate verbose descriptions of each image, embed the descriptions using another model, which I also use for the query embedding.

The stack is hacky, since it was mostly for myself...

alt Hacker News