I'm really looking for a multi-modal image capable version of Jev.
If we could get machine learning type results on images without training, that would be fantastic.
Jev’s context should have signals/features to operate on. Same way LLMs can use CV & code to analyze an image.
Jev’s context should have signals/features to operate on. Same way LLMs can use CV & code to analyze an image.