Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....
Cloudflare’s clef is multimodal (https://blog.cloudflare.com/clef-decision-models/)
(Disclaimer, I work at Cloudflare, but not on models)
Strands Decider can take vision in. How does it go with that question?
Cloudflare’s clef is multimodal (https://blog.cloudflare.com/clef-decision-models/)
(Disclaimer, I work at Cloudflare, but not on models)