> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.
I'm confused, videos contain images and audio ...?
It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.
That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.