gemini flash is probably the best model for visual tasks right now. they also make it really easy to ingest videos
Yes, was going to say I use it exclusively for video and audio. The ability to give it a YouTube link through the API and ask questions about it is awesome
Ah, multimodal is a great point. I'll need to try that some time.
Crazy it's still the only video understanding endpoint. It's what I use it for and no other model even offers a competitor.