logoalt Hacker News

SubiculumCode • today at 5:34 AM • 2 replies • view on HN

Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....


Replies

necubi • today at 5:53 AM

Cloudflare’s clef is multimodal (https://blog.cloudflare.com/clef-decision-models/)

(Disclaimer, I work at Cloudflare, but not on models)

sauhsoj • today at 5:54 AM

Strands Decider can take vision in. How does it go with that question?

➕ show 1 reply