Finally. I was getting annoyed that there's been an inflection point in how LLMs/agents work but there hasn't been a good moderate-size embeddings model, and this one is multimodal too! 270M for text only is great compared to older embedding models, and a total 440M for text + vision is also fair.
I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.
Not just vision with video, but also audio, it really seems amazing.
I’m not sure how it can handle vicinity of pairs of embeddings with for example some words and the audio where they’re spoken or an image where the text is handled. Building local multimodal search with this would be amazing. I’ve explored this stuff with CLIP and it’s interesting how image (but also audio) embedding carries both the clean “text” content information but also the stylistic and visual/audio tone information, the two can even kind of be linearly separated.